Control device, control system, control method, and computer-readable medium

CN117621042BActive Publication Date: 2026-09-22TOYOTA JIDOSHA KK
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
CN202311083329.4
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Priority Date
2022-08-29
Filing Date
2023-08-25
Publication Date
2026-09-22
Estimated Expiration
2043-08-25

AI Technical Summary

Benefits of technology

[0022]根据本公开,能够提供一种即使在任务未知的环境下,也能够有效地实现任务的目标的控制装置、控制系统、控制方法以及程序。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN117621042B_ABST
    Figure CN117621042B_ABST
Patent Text Reader

Abstract

The present disclosure provides a control device, a control system, a control method, and a computer readable medium. In the present disclosure, a request response processing section calculates a request parameter related to whether to request assistance and a response parameter related to whether to respond to a request from another agent based on observation information related to the agent, another agent in the vicinity of the agent, and a task. An importance degree processing section performs processing for calculating an importance degree of each of the tasks related to the agent based on at least the request parameter of another agent and the response parameter of the agent. A task selection section selects a task that the agent should perform according to the importance degree. A task execution section controls the agent so as to perform the selected task.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This disclosure relates to a control device, a control system, a control method, and a program. Background Technology

[0002] There exists a technology that enables multiple intelligent agents (such as robots) to perform tasks. Related to this technology, Japanese Patent Application Publication No. 2017-094122 discloses a mobile intelligent agent capable of assembling general-purpose structures. In Japanese Patent Application Publication No. 2017-094122, multiple mobile intelligent agents automatically manipulate components such as modules on a work surface to perform actions like assembling general-purpose structures. Furthermore, various mobile intelligent agents sometimes operate in a cooperative manner. Summary of the Invention

[0003] In environments where the task is unknown, the number of agents required to perform the task may be unpredictable. In such cases, as described in Japanese Patent Application Publication No. 2017-094122, where multiple agents work collaboratively to perform the task, the task may fail. Therefore, in the technology disclosed in Japanese Patent Application Publication No. 2017-094122, the task objective may not be effectively achieved.

[0004] This disclosure provides a control device, control system, control method, and program that can effectively achieve the objectives of a task even in environments where the task is unknown.

[0005] The control device disclosed herein is a control device for controlling an intelligent agent performing a task, wherein, for the task, the more intelligent agents performing the task, the higher the probability of achieving the goal of the task, and for the task, there are multiple intelligent agents in the environment. The control device comprises: a request response processing unit, which calculates request parameters related to whether to request support and response parameters related to whether to respond to requests from other intelligent agents based on observation information related to the intelligent agent, other intelligent agents around the intelligent agent, and the task; an importance processing unit, which implements processing for calculating the importance of each of the tasks related to the intelligent agent based at least on the request parameters of other intelligent agents and the response parameters of the intelligent agent; a task selection unit, which selects the task that the intelligent agent should perform according to the importance; and a task execution unit, which controls the intelligent agent to perform the selected task.

[0006] Furthermore, the control system disclosed herein is a system for decentralized control of multiple intelligent agents performing a task, wherein, for the task, the more intelligent agents performing the task, the higher the probability of achieving the task's objective, and for the task, multiple agents exist in the environment. The control system has multiple control devices for controlling the multiple intelligent agents respectively. Each control device has: a request response processing unit, which calculates request parameters related to whether to request support and response parameters related to whether to respond to requests from other intelligent agents based on observation information related to the intelligent agent of the control device, other intelligent agents around the intelligent agent, and the task; an importance processing unit, which implements processing for calculating the importance of each task related to the intelligent agent based at least on the request parameters of other intelligent agents and the response parameters of the intelligent agent; a task selection unit, which selects the task that the intelligent agent should perform according to the importance; and a task execution unit, which controls the intelligent agent to perform the selected task.

[0007] Furthermore, the control method disclosed herein is a control method for controlling an intelligent agent performing a task. For the task, the more intelligent agents performing the task, the higher the probability of achieving the task's objective. For the task, multiple agents exist in the environment. In the control method, based on observation information related to the intelligent agent, other intelligent agents in the vicinity of the intelligent agent, and the task, request parameters related to whether to request support and response parameters related to whether to respond to requests from other intelligent agents are calculated. A process is implemented to calculate the importance of each task related to the intelligent agent based at least on the request parameters of other intelligent agents and the response parameters of the intelligent agent. Based on the importance, the task that the intelligent agent should perform is selected, and the intelligent agent is controlled to perform the selected task.

[0008] Furthermore, the program involved in this disclosure is a program for implementing a control method for controlling an intelligent agent performing a task, wherein, for the task, the more intelligent agents performing the task, the higher the probability of achieving the goal of the task, and for the task, there are multiple agents in the environment. The program causes a computer to perform the following steps: calculating request parameters related to whether to request support and response parameters related to whether to respond to requests from other intelligent agents based on observation information related to the intelligent agent, other intelligent agents around the intelligent agent, and the task; performing a process for calculating the importance of each of the tasks related to the intelligent agent based at least on the request parameters of other intelligent agents and the response parameters of the intelligent agent; selecting the task that the intelligent agent should perform according to the importance; and controlling the intelligent agent to perform the selected task.

[0009] In this disclosure, the objectives of the mission can be effectively achieved even in environments where the mission is unknown.

[0010] Furthermore, preferably, the request response processing unit calculates the request parameters and the response parameters based on a strategy learned for each of the agents.

[0011] In this disclosure, such a structure enables the appropriate selection of tasks to be performed for each agent.

[0012] Furthermore, preferably, the request response processing unit calculates the request parameters and the response parameters based on the request level and response level that are input into the strategy and output from the strategy, respectively.

[0013] In this disclosure, such a structure enables the appropriate selection of tasks to be performed for each agent.

[0014] Furthermore, preferably, the request response processing unit calculates the request parameters indicating a request for support when the request level exceeds a predefined threshold and the task that the agent is performing or wants to perform has not been performed.

[0015] In this disclosure, through such a structure, it is possible to appropriately calculate the request parameters indicating the need for support when support should be requested for a task that the agent is performing or wants to perform.

[0016] Furthermore, preferably, the request response processing unit calculates the response parameters indicating the status of the response request when the response level exceeds a predefined threshold and the task that the agent is performing or wants to perform has not been performed.

[0017] In this disclosure, such a structure enables the intelligent agent to continue performing a task while it is performing or wants to perform the task.

[0018] Furthermore, preferably, the importance processing unit calculates the importance of each of the tasks associated with each agent based on the strategy learned for each agent.

[0019] In this disclosure, such a structure enables the appropriate calculation of the importance of each task for each agent.

[0020] Furthermore, preferably, the importance processing unit calculates the importance of the task corresponding to the observation information associated with the agent based on a target value of the importance of the task corresponding to the observation information, which is input into the strategy and output from the strategy.

[0021] In this disclosure, a structure allows for the calculation of the task importance corresponding to the observed information for each agent in a manner close to the target value. This enables the appropriate calculation of task importance.

[0022] According to this disclosure, a control device, control system, control method, and program can be provided to effectively achieve the objectives of a task even in environments where the task is unknown.

[0023] The above and other objects, features and advantages of this disclosure will become more fully understood from the detailed description and accompanying drawings given below, which are given by way of illustration only and should not be considered as limiting the disclosure. Attached Figure Description

[0024] Figure 1 A diagram illustrating the control system involved in Implementation Method 1.

[0025] Figure 2 This diagram illustrates the structure of the control device according to Embodiment 1.

[0026] Figure 3 This diagram illustrates the environment in which the intelligent agent and the task described in Implementation 1 exist.

[0027] Figure 4This is a flowchart illustrating the control method executed by the control device according to Embodiment 1.

[0028] Figure 5 This is a flowchart illustrating the control method executed by the control device according to Embodiment 2.

[0029] Figure 6 This is a flowchart illustrating the control method executed by the control device according to Embodiment 3. Detailed Implementation

[0030] (Implementation Method 1)

[0031] Hereinafter, this embodiment will be described with reference to the accompanying drawings. For clarity of explanation, the following description and drawings have been appropriately omitted and simplified. Furthermore, in the various drawings, the same symbols are used for the same elements, and repeated descriptions have been omitted as needed.

[0032] Figure 1 Figure 1 illustrates the control system 1 according to Embodiment 1. The control system 1 includes a control device 100 that controls multiple agents 10 respectively, and a monitoring device 60 that monitors multiple tasks 50 respectively. While the agents 10 may be machines such as robots, they are not limited thereto. Each agent 10 is configured in an environment and autonomously performs actions within the environment through control implemented by the control device 100.

[0033] The control device 100 is, for example, a computer. The control device 100 may also be built into an intelligent agent 10, such as a robot. The control device 100 implements control to cause the corresponding intelligent agent 10 to perform task 50. That is, the control system 1 provides distributed control over multiple intelligent agents 10. Each control device 100 is communicatively connected to other control devices 100 via a wired or wireless network. Furthermore, each control device 100 is communicatively connected to a monitoring device 60 via a wired or wireless network. Details regarding the control device 100 will be described later.

[0034] In the environment where agent 10 exists, there are multiple tasks 50. Agent 10 executes each of the multiple tasks 50. In each task 50, a goal (end point; termination condition) is set. Tasks 50 are made to proceed by each agent 10 executing each task 50, and each task 50 is completed (terminated) by achieving its goal.

[0035] Here, for task 50, the more agents 10 that execute task 50, the higher the feasibility (the probability of achieving the goal of task 50). That is, even if one agent 10 attempts to execute a task 50 but the task 50 is not performed, executing the task 50 by multiple agents 10 increases the probability that task 50 can be performed and that its goal (the probability of achieving the goal of task 50) is higher. In other words, coordinating the execution of task 50 by multiple agents 10 increases the probability of task 50 being achieved. In other words, coordinating the execution of task 50 by multiple agents 10 increases the probability that the goal of task 50 is achieved. However, the number of agents 10 required for task 50 to proceed may not be known beforehand. The number of agents 10 required for task 50 to proceed can be determined by the agents 10 executing task 50. The control device 100 controls the agents 10 in a manner that task 50 is executed to achieve its goal. Details will be described later.

[0036] The monitoring device 60 is, for example, a sensor or a camera. The monitoring device 60 monitors (detects) the status of each task 50. Specifically, the monitoring device 60 detects, for example, the position and speed of the task 50. Furthermore, the monitoring device 60 stores information on whether the task has ended. Additionally, the monitoring device 60 may also store information related to the objective of the task 50. The monitoring device 60 may also monitor whether the objective of the task 50 has been achieved. Furthermore, the monitoring device 60 may be set up for each task 50. Alternatively, one monitoring device 60 may monitor multiple tasks 50. Additionally, each intelligent agent 10 may also detect the status of the task 50. In this case, the monitoring device 60 may not be present. Furthermore, the intelligent agent 10 may also detect the status of the task 50 and determine the end of the task 50.

[0037] In this embodiment 1, task 50 refers to goods that need to be moved. Furthermore, in each task 50, a destination (target) is set as the destination for the goods. Intelligent agents 10 move the goods (task 50) to the destination in a manner that allows the goods (task 50) to reach the destination. Moreover, the more intelligent agents 10 that move the goods (task 50), the higher the probability that the goods (task 50) will reach the destination. That is, depending on the goods, the goods may be too large to move for a small number of intelligent agents 10. In other words, the size and weight may vary depending on the goods. On the other hand, by coordinating the movement of a large number of intelligent agents 10, the goods can be moved. That is, by cooperating with a large number of intelligent agents 10, the goods can be moved (task 50 is performed) and moved to their target location (achieving the target of task 50). Furthermore, the number of intelligent agents 10 required to move the goods is unknown. The number of intelligent agents 10 required for the movement can only be determined after the intelligent agents 10 attempt to perform the movement.

[0038] Figure 2 This is a diagram illustrating the structure of the control device 100 according to Embodiment 1. Figure 2 As shown, the control device 100 includes a control unit 102, a storage unit 104, a communication unit 106, and an interface unit 108 (IF). The control unit 102, storage unit 104, communication unit 106, and interface unit 108 are interconnected via a data bus or the like. Additionally, the machine's intelligent agent 10 may also have... Figure 2 The hardware structure of the control device 100 is shown. Furthermore, the monitoring device 60 may also have... Figure 2 The hardware structure of the control device 100 shown.

[0039] The control unit 102 is, for example, a processor such as a CPU (Central Processing Unit). The control unit 102 functions as a computing device that performs control processing and arithmetic processing. Furthermore, the control unit 102 may have multiple processors. The storage unit 104 is, for example, a storage device such as a memory or a hard disk. The storage unit 104 is, for example, a ROM (Read Only Memory) or RAM (Random Access Memory). The storage unit 104 has the function of storing control programs and arithmetic programs executed by the control unit 102. That is, the storage unit 104 (memory) stores more than one command. In addition, the storage unit 104 has the function of temporarily storing processing data. The storage unit 104 may contain a database. Furthermore, the storage unit 104 may have multiple memories.

[0040] The communication unit 106 performs the processing necessary for communication with other devices such as the control device 100 or monitoring device 60 via a network. The communication unit 106 may include a communication port, router, firewall, etc. The interface unit 108 is, for example, a user interface (UI). The interface unit 108 has input devices such as a keyboard, touch panel, or mouse, and output devices such as a display or speaker. The interface unit 108 may also be configured, for example, like a touchscreen (touch panel), to integrate the input and output devices. The interface unit 108 accepts data input operations performed by the user (operator) and outputs information to the user.

[0041] In the control device 100 according to Embodiment 1, structural elements include an observation information acquisition unit 110, a strategy storage unit 112, an action output unit 120, a request response processing unit 130, an importance processing unit 140, a task selection unit 150, and a task execution unit 160. Each of these structural elements can be implemented, for example, by executing a program under the control of the control unit 102. More specifically, each structural element can be implemented by the control unit 102 executing a program (command) stored in the storage unit 104. Alternatively, each structural element can be implemented by pre-recording the required program in any non-volatile recording medium and installing it as needed. Furthermore, each structural element is not limited to implementation by program-based software; it can also be implemented by any combination of hardware, firmware, and software. Additionally, each structural element can be implemented, for example, using a user-programmable integrated circuit such as a FPGA (field-programmable gate array) or a microcomputer. In this case, the integrated circuit can also be used to implement the program composed of the aforementioned structural elements. These conditions also apply to other implementations described later.

[0042] Furthermore, in the following description, the control device 100 to be described will be referred to as "its own control device 100 (the control device)". Additionally, control devices 100 other than its own control device 100 will be referred to as "other control devices 100". Furthermore, the intelligent agent 10 controlled by its own control device 100 will be referred to as "its own intelligent agent (the intelligent agent)". Furthermore, intelligent agents 10 other than its own intelligent agent 10 will be referred to as "other intelligent agents". Furthermore, in the following description, the operation of its own control device 100 will be described, but the same operation will be performed for other control devices 100.

[0043] The control device 100 controls its own intelligent agent 10 through the aforementioned structural elements to execute task 50 in a manner that achieves the objective of task 50. That is, the control device 100 implements control for its own intelligent agent 10 to execute task 50. The control device 100 calculates request parameters and response parameters for its own intelligent agent 10 based on its own intelligent agent 10, other intelligent agents 10 in its vicinity, and observation information related to task 50. Here, the "request parameter" is a parameter related to whether to request support from other intelligent agents 10. Furthermore, "requesting support" corresponds to the situation where other intelligent agents 10 perform task 50 in a manner that coordinates with its own intelligent agent 10. Furthermore, the "response parameter" is a parameter related to whether to respond to requests from other intelligent agents 10. Furthermore, "responding to a request" corresponds to the situation where its own intelligent agent 10 performs task 50 in a manner that coordinates with other intelligent agents 10.

[0044] Furthermore, the control device 100 performs a process to calculate the importance of each task 50 related to its own agent 10 based on request parameters from other agents 10 and response parameters from its own agent 10. Here, "importance" is used to determine which task 50 the agent 10 selects and executes. The higher the importance of a task 50, the higher the likelihood that it will be selected by the agent 10 and executed by the selected agent 10.

[0045] Furthermore, the control device 100 selects the task 50 that its own agent 10 should perform based on its importance. The control device 100 controls the execution of the selected task 50 by its own agent 10. Then, the control device 100 repeatedly performs the above-described processing for each control cycle. The likelihood of the importance of the task 50 performed by the agent 10 that has calculated the request parameters indicating a request for support being greater than that of other agents 10 that have calculated the response parameters indicating a response to the request is higher. Therefore, the likelihood of that agent 10 supporting the task 50 is higher. This will be explained in detail below.

[0046] The observation information acquisition unit 110 acquires observation information from the surrounding environment. This observation information includes information related to its own agent 10, other agents 10 in the vicinity of its own agent 10, and the task 50. Therefore, the observation information includes information related to its own agent 10. Furthermore, the observation information includes information related to its own agent 10's vicinity, information related to other agents 10, and information related to the task 50.

[0047] Figure 3This diagram illustrates the environment in which the agent 10 and task 50 according to Embodiment 1 exist. The number of agents 10 is set to M, and the number of tasks 50 is set to N. Furthermore, the agent 10 itself is designated as "Agent #i", where i is the index representing the agent 10. Other agents 10 are designated as "Agent #j", where j is the index representing the other agent 10.

[0048] Furthermore, other agents 10 near agent #i are referred to as "nearby agents". Nearby agents can also be those whose distance from agent #i is within a predetermined range (e.g., ...). Figure 3 The predetermined number of other agents 10 are represented by dashed circles. Alternatively, the nearby agents can be the predetermined number of other agents 10 closest to agent #i. In embodiment 1, the "predetermined number" is set to 2. These two nearby agents are designated as agent #j1 and agent #j2, respectively. Furthermore, in Figure 3 The diagram shows agent #1 and agent #M, which are agents 10 other than nearby agents. Additionally, there are actually (M-3) agents 10 other than nearby agents, obtained by removing the agent 10 itself and two nearby agents from the total number of agents 10 M.

[0049] Furthermore, the index of task 50 is set to "l" (l∈{1, ..., N}). Tasks 50 near agent #i are called "nearby tasks". A nearby task can also be, for example, a task whose distance from agent #i is within a predetermined range (e.g., ..., N). Figure 3 The predetermined number of tasks 50 is represented by a dashed circle. Alternatively, this "predetermined range" can be a range different from the range defined for the nearby agents described above. Or, the nearby tasks can be the predetermined number of tasks 50 closest to agent #i. In Embodiment 1, the "predetermined number" is set to 2. These two nearby tasks are designated as task #11 and task #12, respectively. Furthermore, in Figure 3 The diagram shows tasks 50 other than nearby tasks, namely tasks #1, #2, and #N. Additionally, there are actually (N-2) tasks 50 other than nearby tasks, obtained by removing nearby tasks from the total number N of tasks 50.

[0050] Furthermore, x represents the position (current position) of agent 10. i Indicates the intelligent agent # i The location. x j Indicates the intelligent agent # j The position. Additionally, z represents the position of task 50 (current position). Furthermore, z * This indicates the target location (end point) for Task 50.l Indicates the position of task #l. l * This indicates the target location of task #1. Furthermore, the "location" of task 50 is not limited to the physical location where task 50 exists; it can also refer to the state of task 50. In this case, the "location" of task 50 can also represent a point in virtual space that reflects the state of task 50. For example, the state of task 50 can also represent the progress of task 50, and the "location" of task 50 can also represent a point in virtual space that reflects the progress of task 50.

[0051] Furthermore, φ represents the importance of each task 50 associated with each agent 10. i φ represents the importance of each task 50 for agent #i. j This represents the importance of each task 50 for agent #j. Additionally, φ has a number of components corresponding to the number N of tasks 50, and shows the importance of each of tasks #1, ..., #l, ..., #N. For example, the importance φ... i l This indicates the importance of task #l to agent #i. The importance associated with each agent 10 is calculated in the control device 100 of each agent 10 and transmitted (broadcast) to the surrounding agents 10 (control devices 100). Details will be described later.

[0052] The observation information acquisition unit 110 acquires the positions of the surrounding agents 10 and tasks 50. Specifically, the observation information acquisition unit 110 acquires information related to the agent 10 from other control devices 100 associated with the surrounding agent 10. This information includes, for example, the position of the agent 10 and the importance of each task 50 associated with the agent 10 (the importance of each task 50 to the agent 10). Furthermore, the observation information acquisition unit 110 acquires information related to each task 50 from the monitoring device 60. This information includes, for example, the state of the task 50 and the goal of the task 50. The state of the task 50 may include, for example, the position and speed of the task 50.

[0053] The observation information acquisition unit 110 calculates the distance D between agent #i and agent #j based on the acquired position of agent #j. ij Perform the calculation. Here, D... ij =||x i -x j2. Furthermore, the observation information acquisition unit 110 calculates the distance between its own agent 10 (agent #i) and each task #l based on the acquired position of the task 50. Specifically, the observation information acquisition unit 110 uses the following formula (1) to calculate the distance D between agent #i and task #l. il Perform the calculation.

[0054] [Mathematical Expression 1]

[0055]

[0056] In equation (1), "0.05" is the threshold used to determine whether task #l has reached the target position (i.e., whether task #l has achieved the target). If z l With z l * If the distance between them is less than 0.05, then task #1 can be considered to have reached the target position. Furthermore, "1.0e4" is a value large enough that task #1 is not considered to be near agent #i. That is, according to equation (1), for task 50 that has reached the target position, D... il The distance was calculated to be very large compared to the actual distance. Therefore, for task 50 that has reached the target location, it can be ignored in subsequent processing.

[0057] The observation information acquisition unit 110 uses D ij and D il To obtain observation information from its own agent 10 (agent #i) i The observation information acquisition unit 110 uses D... ij To determine a predetermined number of nearby agents, and to set information related to nearby agents as observation information. i Part of it. In addition, the observation information acquisition unit 110 uses D... il To determine a predetermined number of nearby tasks, and to set information related to nearby tasks as observation information. i Part of it.

[0058] Here, the nearby tasks in Implementation 1 will be explained. The conditions of the nearby tasks related to agent #i are expressed by the following equation (2). In addition, as shown in equation (2), the number of nearby tasks related to agent #i is two.

[0059] [Mathematical Expression 2]

[0060]

[0061] Here, l i 1Let #l1 be the nearby task #l1 associated with agent #i, and defined by the following equation (3). That is, the nearby task #l i 1 (Nearby Task #l1) refers to the task 50 closest to agent #i. Additionally, this nearby task #l1 could be task 50 that agent #i is currently executing at the current time.

[0062] [Mathematical Expression 3]

[0063]

[0064] In addition, l i 2 Let #l2 be the nearby task #l2 associated with agent #i, and defined by the following equation (4). That is, the nearby task #l i 2 (Nearby task #l2) is task 50, which is the second closest to agent #i.

[0065] [Mathematical Expression 4]

[0066]

[0067] The observation information acquisition unit 110 acquires observation information as shown in the following equation (5). i Furthermore, in equation (5), the T in the upper right corner represents the transpose. Additionally, equation (5) represents the observation information at a certain point in time (e.g., time t).

[0068] [Mathematical Expression 5]

[0069]

[0070] Here, in equation (5), the following equation (6) represents information about the agent 10 (agent #i). In addition, equation (6) represents, from left to right, the position of agent #i, the importance of nearby task #l1 related to agent #i, and the importance of nearby task #l2 related to agent #i.

[0071] [Mathematical Expression 6]

[0072]

[0073] Furthermore, in equation (5), the following equation (7) represents information about the nearby agent #j1. Similarly, like the nearby task #l1, the nearby agent #j1 can also be any other agent 10 closest to its own agent 10 (agent #i). Additionally, equation (7) from left to right represents the position of the nearby agent #j1, the importance of the nearby task #l1 associated with the nearby agent #j1, and the importance of the nearby task #l2 associated with the nearby agent #j1.

[0074] [Mathematical Expression 7]

[0075]

[0076]

[0077] Furthermore, in equation (5), the following equation (8) represents information about the nearby agent #j2. Similarly, like the nearby task #l2, the nearby agent #j2 can also be another agent 10 that is second closest to its own agent 10 (agent #i). Additionally, equation (8) sequentially represents, from left to right, the position of the nearby agent #j2, the importance of the nearby task #l1 associated with the nearby agent #j2, and the importance of the nearby task #l2 associated with the nearby agent #j2.

[0078] [Mathematical Expression 8]

[0079]

[0080] Furthermore, in equation (5), the second o_(l1)^task from the right represents information related to the nearby task #l1. o_(l1)^task can also represent the state and target of the nearby task #l1. As described above, in embodiment 1, task 50 is the goods that should be transported. In this case, o_(l1)^task can be defined as follows: equation (9). In addition, the right side of equation (9) represents, from left to right, the position of the nearby task #l1, the speed of the nearby task #l1, and the target position (the destination, i.e., the transport destination) of the nearby task #l1. The position and speed of the nearby task #l1 correspond to the state of the nearby task. In addition, the speed of the nearby task can be calculated based on the difference in position of the nearby task in each control cycle. The same applies to the speeds of other objects described later.

[0081] [Mathematical Expression 9]

[0082]

[0083] Similarly, in equation (5), the first o_(l2)^task from the right represents information related to the nearby task #l2. o_(l2)^task can also represent the state and target of the nearby task #l2. Furthermore, in embodiment 1, o_(l2)^task can be defined as shown in equation (10) below. Additionally, the right side of equation (10) sequentially represents, from left to right, the position of the nearby task #l2, the speed of the nearby task #l2, and the target position (destination, i.e., the transport destination) of the nearby task #l2. The position and speed of the nearby task #l2 correspond to the state of the nearby task.

[0084] [Mathematical Expression 10]

[0085]

[0086]

[0087] According to equations (5), (9), and (10), in implementation method 1 where the task is to transport goods, the observation information o i It is expressed as in equation (11) below.

[0088] [Mathematical Expression 11]

[0089]

[0090] The policy storage unit 112 stores the policy π (learned model) that has been learned through reinforcement learning. The policy π is learned for each agent 10. Therefore, the learned policy π (the parameters of the network (neural network, etc.) that constitutes the policy π) can be different for each agent 10.

[0091] The policy π of agent #i NN,i The observation information described above i Let this be the input, and the action a will be shown in the following equation (12). i Set as output. Therefore, action a i For, strategy π NN,i The output value.

[0092] [Mathematical Expression 12]

[0093]

[0094] Here, c i l Let φ be the importance of nearby tasks #l related to agent #i. i l The target value, agent #i, corresponds to an indicator (meaning) representing how much importance is attached to nearby tasks #l. Additionally, the importance φ... il It can be taken up to the target value c. i l The value up to that point. In other words, the importance φ i l It can be increased until the target value c is reached. i l Up to this point. c, as the first component of the first right-hand side (the second expression from the left) of equation (12). i ^(l1) is the target value of the importance of the nearby task #l1 associated with agent #i. Similarly, c is the second component on the first right-hand side of equation (12). i ^(l2) is the target value for the importance of the nearby task #l2 related to agent #i.

[0095] Furthermore, a, as the third component of the first right-hand side of equation (12) i d This represents the degree of request from agent #i. Furthermore, a, as the fourth component on the first right-hand side of equation (12),... i σ This represents the responsiveness of agent #i. Here, as shown in equation (13) below, a i d and a i σ The possible values ​​are 0 or higher and 1 or lower.

[0096] [Mathematical Expression 13]

[0097]

[0098] Request level a i d This indicates the degree to which agent #i requests support. In other words, a i d The higher the value of a, the more likely it is that other agents 10 will coordinate and cooperate with its own agent 10 to perform task 50. In other words, a i d The higher the value, the higher the probability that the request parameters for requesting support from other agents 10 will be calculated.

[0099] Response level a i σ This indicates the degree to which agent #i responds to a request. In other words, a i σ The higher the value, the more likely it is that its own agent 10 will coordinate and cooperate with other agents 10 to perform task 50. In other words, a i σThe higher the value, the higher the likelihood that the response parameters will be calculated in response to requests from other agents 10.

[0100] In addition, strategy π NN,i So that the reward r expressed by the following equation (14) is i (t) is maximized during the learning phase. That is, during the learning phase, the policy π is maximized. NN,i Observation information o i Set as input and output action a. i Then, after outputting action a... i In the case of reward r i The calculation is performed using (t). Then, the policy π is updated continuously in a way that maximizes its reward (cumulative reward). NN,i The parameters (weights) of the network within. Furthermore, since there is no index i on the right-hand side of equation (14), the reward is independent of the agent and is common. Then, based on the obtained common reward, the network with different Q values ​​for each agent and the network with the policy are updated. Thus, the policy π NN,i They were able to learn.

[0101] [Mathematical Expression 14]

[0102]

[0103] t represents time. Furthermore, P l (t) represents the achievement level of task #l at time t. Therefore, the first term on the right side of equation (14) is the achievement level P of each task #l (#1, ..., #N) at time t. l The sum of (t) (Summation: cumulative summation). Alternatively, the completion rate of task 50 can also represent the degree to which the objective of task 50 is achieved. Or, the completion rate of task 50 can also represent whether the objective of task 50 has been achieved.

[0104] Furthermore, Q1(t) represents the progress of task #l at time t. Therefore, the second term on the right-hand side of equation (14) is the sum of the progress Q1(t) of each task #l (#1, ..., #N) at time t. λ is a predetermined coefficient. In addition, the progress of task 50 refers to how much task 50 has progressed. That is, the progress of task 50 refers to the progress of task 50. Therefore, if task 50 is in progress, the progress may be high; if task 50 is delayed, the progress may be low.

[0105] Here, as described above, in Embodiment 1, task 50 is the goods to be transported. Moreover, in Embodiment 1, achieving the goal of task 50 means that task 50 reaches the target location. Therefore, in Embodiment 1, P1(t) is defined as follows (15). For equation (15), if task #1, as the goods, reaches the target location at time t, then P1(t) = 1; otherwise, P1(t) = 0.

[0106] [Mathematical Expression 15]

[0107]

[0108] Furthermore, in Embodiment 1, it can be said that if the goods move quickly, the handling of the goods progresses smoothly, and if the goods do not move quickly, the handling of the goods is delayed. That is, in Embodiment 1, the faster the task 50, which is the goods, moves, the more likely the progress of the task 50 is to increase. Therefore, in Embodiment 1, Q1(t) is defined as shown in Equation (16) below. As shown in Equation (16), in Embodiment 1, the progress Q1(t) of task #1 corresponds to the moving speed of task #1, which is the goods. That is, in Embodiment 1, the faster the task #1, which is the goods, moves, the more likely the progress Q1(t) is to increase.

[0109] [Mathematical Expression 16]

[0110] Q l (t): = ||v l (t)||2

[0111] …(16)

[0112] According to equation (16), in embodiment 1, the reward r represented by equation (14) above is... i (t) is shown as in equation (17) below.

[0113] [Mathematical Expression 17]

[0114]

[0115] Action output unit 120 uses the strategy π described above to output observation information o. i Corresponding action a i Specifically, the action output unit 120 will observe the information. i Input to policy π NN,i Therefore, strategy π NN,i Output action a i .

[0116] The request response processing unit 130 calculates the request parameters and response parameters related to its own intelligent agent 10. The request response processing unit 130 calculates the parameters based on the policy π obtained from the action output unit 120. NN,i Output action a i This allows for the calculation of request and response parameters related to agent #i. Here, due to action a... i It is based on observation information. i The output is, therefore, the request and response processing unit 130 calculates the request and response parameters based on the observation information.

[0117] The request response processing unit 130 is based on the action output unit 120 from the strategy π NN,i Output request level a i d Thus, the request parameter d related to agent #i is... i Calculations are performed. Specifically, the request response processing unit 130 can also perform calculations at request level a. i d If the threshold is exceeded, the request parameter d, which indicates a request for support, will be used. i Calculations are performed. On the other hand, the request response processing unit 130 can also perform calculations at request level a. i d For cases below the threshold, the request parameter d indicates that no support is requested. i Perform calculations. Alternatively, the request response processing unit 130 can also perform calculations at request level a. i d If the threshold is exceeded and the task 50 that the agent 10 is performing or intends to perform has not been performed, the request parameter d, which indicates a request for support, is used. i Calculations are performed. On the other hand, the request response processing unit 130 may also, if the above conditions are not met, process the request parameter d indicating that support is not requested. i The calculation is performed. The request parameter d is calculated by the request response processing unit 130. i Send to the control device 100 associated with other intelligent agents 10.

[0118] For example, the request response processing unit 130 uses the following equation (18) to process the request parameter d of agent #i. i Perform the calculation. Here, d i =1 indicates that agent #i requests support. d i =0 indicates that agent #i does not request support. Therefore, the request parameter d is... i It can function as a trigger for the event that requests support.

[0119] [Mathematical Expression 18]

[0120]

[0121] In equation (18), "0.5" is a predefined threshold. The threshold is not limited to 0.5. Furthermore, l i * This indicates that task 50 is currently selected for agent #i. In other words, l i * This indicates that task 50 was selected in the previous control cycle via task selection unit 150 (described later). In other words, l i * This indicates that its agent 10 (agent #i) is performing or wants to perform task 50. Additionally, Q_(l i * (t) is for task #l i * The progress. Moreover, the so-called Q_(l i * (t) = 0, which indicates that task #l i * The situation that was not carried out.

[0122] Therefore, equation (18) represents the degree of request a i d The threshold "0.5" is exceeded, and for agent #i, the currently selected task is #l. i * The progress is 0 (that is, task #l) i * If not performed, the request parameter is d. i =1. In other words, equation (18) represents the case where the request level a is 1. i d The threshold "0.5" is exceeded, and for agent #i, the currently selected task is #l. i * If no action is taken, calculate the request parameter d, which indicates the request for support. i Furthermore, equation (18) indicates that d... i = 0. That is, equation (18) represents the calculation of the request parameter d, which indicates the case where no support is requested, when the above conditions are not met. i In this embodiment, although request parameters representing a request for support are calculated, in reality, the currently selected task #l for agent #i is not considered. i * In other words, other intelligent agents #j may not necessarily come to provide support. Regarding task #li * As for whether other intelligent agents will actually come to provide support, it depends on the task. i * The importance is determined by the task #l. i * Whether other intelligent agents #j will actually come to provide support will depend on the tasks #l that intelligent agent #j has. i * The importance is determined by the task. In other words, for task #l i * Whether other intelligent agents #j will actually come to provide support will depend on the task #l for agent #j. i * The importance is determined by [the degree of importance].

[0123] Additionally, regarding the task #l selected for its own intelligent agent #i i * Even if one's own agent #i does not cooperate with other agents #j while the task is in progress, i * This will also happen. In such cases, support will be requested, and tasks will be performed in a coordinated manner with other intelligent agents. i * The situation may become useless. Therefore, in equation (18) above, even if the request level a i d The higher the value, the better, for agent #i in relation to the currently selected task #l i * If the progress is not 0, it will also become d. i =0. Therefore, it is possible to suppress the occurrence of useless requests.

[0124] Furthermore, the request response processing unit 130 is based on the action output unit 120 from the strategy π NN,i Output response level a i σ Thus, the response parameter σ associated with agent #i i Calculations are performed. The request-response processing unit 130 can also perform calculations at response level a. i σ If the response parameter σ, representing the response request, is exceeded, the response will be affected if the threshold is exceeded. i Calculations are performed. On the other hand, the request response processing unit 130 can also perform calculations at response level a. i σ When the threshold is below a certain value, the response parameter σ represents the condition of not responding to the request. iPerform calculations. Alternatively, the request response processing unit 130 can also perform calculations at response level a. i σ If the threshold is exceeded and the task 50 that the agent 10 is performing or intends to perform has not been performed, the response parameter σ represents the response request. i Calculations are performed. On the other hand, the request response processing unit 130 may also, if the above conditions are not met, process the response parameter σ indicating a non-response to the request. i Perform the calculation.

[0125] For example, the request response processing unit 130 uses the following equation (19) to process the response parameter σ of agent #i. i Perform the calculation. Here, σ i =1 indicates that agent #i responds to the request. σ i =0 indicates that agent #i does not respond to the request. Therefore, the response parameter σ i It can function as a trigger for the event that responds to a request.

[0126] [Mathematical Expression 19]

[0127]

[0128] In equation (19), "0.5" is a predefined threshold. The threshold is not limited to 0.5. Furthermore, this threshold may not be the same value as the threshold in equation (18). Additionally, as mentioned above, l i * This indicates that task 50 is currently selected for agent #i. Additionally, Q_(l i * (t) is for task #l i * The progress. Moreover, the so-called Q_(l i * (t) = 0, which indicates that task #l i * The situation that was not carried out.

[0129] Therefore, equation (19) represents the response level a i σ The threshold "0.5" is exceeded, and for agent #i, the currently selected task #l i * The progress is 0 (i.e., task #l) i * Without (this step), the response parameter is σ. i =1. In other words, equation (19) represents the case where the response level a = 1. i σThe threshold "0.5" is exceeded, and for agent #i, the currently selected task #l i * If no action is taken, the response parameter σ represents the response to the request. i The calculation is performed under the following conditions. Furthermore, equation (19) indicates the condition where σ is not satisfied. i = 0. That is, equation (19) represents the response parameter σ for the case where the above conditions are not met, indicating that the request will not be responded to. i The calculations are performed. Furthermore, in this embodiment, although response parameters representing the response to the request are calculated, in reality, agent #i may not necessarily support the task #l of other agents #j. Whether agent #i actually supports the task #l of other agents #j depends on the importance of its task #l. In other words, whether agent #i actually supports the task #l of other agents #j depends on the importance of the task #l possessed by agent #j.

[0130] Additionally, if the task #l is selected for one's own agent #i i * If, while in progress, an agent responds to a request to actually support another agent's task #l, then agent #i will stop its task #l. i * The execution of the task #l is possible. However, the situation where agent #i stops the execution of the ongoing task may become useless. That is, the task #l currently selected by agent #i (that is, the task that agent #i is currently executing) may become useless. i * In the case of ongoing tasks, it is preferable to continue executing them. i * Therefore, in equation (19) above, even if the response level a i σ The higher the value, the better, for agent #i in relation to the currently selected task #l i * If the progress is not zero, it also becomes σ. i =0. Therefore, useless responses can be suppressed.

[0131] Here, as described above, in Embodiment 1, task #1 is the goods to be transported. Furthermore, in Embodiment 1, as shown in equation (16) above, the progress of task #1 corresponds to the moving speed of task #1 as the goods. Therefore, in Embodiment 1, task #1... i * The progress is expressed as shown in equation (20).

[0132] [Mathematical Expression 20]

[0133]

[0134] Therefore, in Implementation 1, the request parameter d shown in Equation (18) above is used. i It is expressed as in the following formula (21).

[0135] [Mathematical Expression 21]

[0136]

[0137] Furthermore, in Embodiment 1, the response parameter σ shown in Equation (19) above... i It is expressed as in the following formula (22).

[0138] [Mathematical Expression 22]

[0139]

[0140] The importance processing unit 140 updates (calculates) the importance of each task #l (l∈1,…,N) surrounding its own agent #i. Specifically, the importance processing unit 140 implements a process for calculating the importance of each task related to its own agent #i based on the request parameters of other agents #j and the response parameters of its own agent #i.

[0141] Specifically, the importance processing unit 140 obtains the request parameters d related to agent #j from the control device 100 of the surrounding agent #j. j Furthermore, the importance processing unit 140 obtains the importance φ of each task #l related to each agent #j from the control devices 100 of each surrounding agent #j. j l Furthermore, as described above, the observation information acquisition unit 110 obtains the importance of nearby tasks #l1 and #l2 related to nearby agents #j1 and #j2. On the other hand, the importance processing unit 140 obtains the importance φ of each task #l related to each agent #j not only from nearby agents but also from the control devices 100 of all (acquireable) agents #j in the vicinity.j l Additionally, the importance processing unit 140 can also handle situations where the requested parameter d cannot be obtained from the control device 100 of the surrounding intelligent agent #j due to reasons such as communication failure. j and importance φ j l In this case, d is set for the agent #j. j =0, φ j l =0.

[0142] Furthermore, in the importance processing unit 140, importance φ is used for task #l. i l Importance φ j l Request parameters d related to other smart agents #j j The response parameter σ related to its own agent #i i To determine the importance φ i l Updates are performed. Furthermore, when task #l is a nearby task, the importance processing unit 140 further utilizes the importance φ of nearby tasks #l related to its own agent #i. i l Target value c i l To determine the importance of task #l φ i l Update accordingly. Additionally, the importance φ... i l The importance of the task #l related to one's own intelligent agent #i.

[0143] Specifically, the importance processing unit 140 uses the following formula (23) to assign importance φ to the task #l related to agent #i. i l The change (and the corresponding change) is calculated. Furthermore, in equation (23), k is a predefined coefficient. Additionally, for ease of notation, the left side of equation (23) is sometimes represented as "φ". i l (point)".

[0144] [Mathematical Expression 23]

[0145]

[0146] Importance processing unit 140 processes the current importance φ i l Add the importance φ shown in equation (23) to the top. i l Change φi l (points), thus determining the importance φ of task #l related to agent #i. i l The importance processing unit 140 updates the importance φ of the task #l related to agent #i using the following formula (24). i l Updates are performed. Additionally, Δt represents the control period. The importance processing unit 140 updates the importance of all tasks #l. Furthermore, let φ be the initial value of the importance of all agents #i (i = 1, ..., M) and all tasks #l (l = 1, ..., N). i l (0) has been predefined.

[0147] [Mathematical Expression 24]

[0148]

[0149] In the case where task #l is not a nearby task (task #l does not satisfy the nearby task condition shown in equation (2)), the importance φ i l Change φ i l (Point) is represented as shown in the following expression on the right-hand side of equation (23). The following expression on the right-hand side of equation (23) corresponds to the following value, namely, the request parameter d related to agent #j. j From the importance φ j l Subtract importance φ i l The resulting value, the product of the product with coefficient k, and the sum of all agents #j, are then multiplied by the response parameter σ associated with agent #i. i The value obtained later. Additionally, for the following expression on the right-hand side of equation (23), if the response parameter σ of one's own agent #i... i If it is 0, then it becomes 0. That is, if task #l is not a nearby task, and the response parameter σ... i If the value is 0, then the importance φ of the task #l related to the agent #i is... i l It will not be updated (no change). Furthermore, in the response parameter σ... i When the value is 1, for other agents #j (the requesting agent) with a request parameter of 1, from the perspective of importance φ j l Subtract importance φ i l The sum of the differences obtained corresponds to the change φ. i l(Point). Therefore, the greater the importance φ, the better. j l Specific importance φ i l A much larger request agent #j, or more agents with greater importance φ. j l Greater than importance φ i l If the requesting agent #j is given, then the importance φ of the task #l related to agent #i is given. j l It may become larger.

[0150] On the other hand, when task #l is a nearby task (where task #l satisfies the nearby task conditions shown in equations (2) to (4), the importance φ i l Change φ i l (Point) is represented as shown in the right-hand side of equation (23). The right-hand side of equation (23) corresponds to the following value, namely, the target value c of the importance of task #l related to agent #i, added to the following equation. i l Subtract importance φ i l The value obtained is the product of the obtained value and the coefficient k. Furthermore, the second term of the upper equation on the right-hand side of equation (23) is the same as the lower equation on the right-hand side of equation (23). Therefore, for the second term of the upper equation on the right-hand side of equation (23), if the response parameter σ of one's own agent #i... i If it is 0, then it becomes 0. That is, if task #l is a nearby task, and the response parameter σ... i If the value is 0, then according to the first term of the above equation on the right side of equation (23), the importance φ of the task #l related to one's own agent #i is... i l It can be updated to a value close to the target value c. i l Furthermore, in the response parameter σ i When the value is 1, the importance φ will be determined from the perspective of the requesting agent. j l Subtract importance φ i l The sum of the differences obtained later, and the difference from the target value c i l Subtract importance φ i l The value obtained by adding the product of the obtained value and the coefficient k corresponds to the change φ. i l (Point). Therefore, at the target value c il In cases where the value is large, the greater the importance φ, the better. j l Specific importance φ i l A much larger request agent #j, or one with more importance φ. j l Greater than importance φ i l If the requesting agent #j is given, then the importance φ of the task #l related to agent #i is given. j l It may become larger.

[0151] Furthermore, the importance processing unit 140 performs processing on the importance of task #1, which has achieved its objective. Specifically, the importance processing unit 140 assigns the importance φ of task #1, which has achieved its objective, to... i l Set to 0. Here, as described above, in Embodiment 1, achieving the goal of task #1 means that task #1, as a cargo, reaches the target location. Therefore, the importance processing unit 140 assigns the importance φ of task #1, which has reached the target location, to the following formula (25). i l Set to 0. Additionally, δ is a threshold used to determine whether task #l has reached the target position (i.e., whether task #l has achieved its target), and in examples such as equation (1), it is 0.05. Therefore, processing of task #l that has achieved its target is no longer performed, and each agent 10 will execute other tasks 50.

[0152] [Mathematical Expression 25]

[0153]

[0154] Furthermore, the importance processing unit 140 will calculate the importance φ of each task #l related to agent #i. i l The control device 100 sends the information to other agents #j. That is, the importance of each task 50 related to each agent 10 is shared among the agents 10 (control devices 100). Therefore, the control devices 100 of other agents #j perform the above-described processing for that agent #j. In other words, the control devices 100 of other agents #j assign importance φ to each task #l related to that agent #j. j l Perform calculations (update).

[0155] Task selection unit 150 determines the task that its agent #i should perform. i *The selection is made. Specifically, the task selection unit 150 selects all tasks #l (l=1,…,N) based on the following formula (26), where the importance φ is different. i l The biggest task is choosing the task that should be performed by the agent. i * In Implementation 1, since Task 50 involves goods that should be transported, the Task Selection Unit 150 selects the goods (Task #1) that are most important to its own agent #i. i * Choose from.

[0156] [Mathematical Expression 26]

[0157]

[0158] The task execution unit 160 performs processing for its own intelligent agent #i to execute task #l. Specifically, the task execution unit 160 causes its own intelligent agent #i to execute the task #l selected by the task selection unit 150. i * Control is achieved through this method. More specifically, Task Execution Unit 160 obtains Task #l i * The location and the objective to be achieved (termination condition). Task execution unit 160 moves agent #i to task #l. i * The position. At this time, the task execution unit 160 can also calculate the speed command value of agent #i. Then, the task execution unit 160 executes task #l. i * And accomplish task #l i * The task execution unit 160 controls the arm of the intelligent agent #i in a way that achieves the target objective. When task #l involves cargo, the task execution unit 160 controls the arm of the intelligent agent #i by holding the cargo. At this time, the task execution unit 160 can also calculate the force and torque command values ​​at the tip of the arm (end effector, etc.). The task execution unit 160 then moves task #l... i * Move to task #1 i * The agent #i is controlled by targeting the target position.

[0159] Additionally, in the request parameter d of its own agent #i i When the value is 0, in the processing of the control device 100 of other intelligent agents #j, in the second term of the upper equation on the right side of equation (23) and the lower equation on the right side, it becomes d. j k(φ jl -φ i l = 0. However, since it is the processing in the control device 100 of other intelligent agents #j, please note that j corresponds to its own intelligent agent #i, and i corresponds to other intelligent agents #j. Therefore, in the request parameter d of its own intelligent agent #i i When the value is 0, in the processing of the control device 100 of other agents #j, even if the importance of task #l is high in one's own agent #i, it is highly likely that the importance of tasks #l related to other agents #j will not be affected. Here, the task #l that is high in one's own agent #i may include the task #l selected for one's own agent #i. i * Based on the above, in the request parameter d i When the value is 0, in the processing of the control device 100 of other intelligent agents #j, the task #l selected for its own intelligent agent #i is chosen. i * The likelihood of this happening decreases. Therefore, other intelligent agents #j come to support the task #l i * The likelihood of that decreases.

[0160] In contrast, the request parameter d of its own agent #i i When the value is 1, in the processing of the control device 100 of other agents #j, the likelihood that the importance (weight) of task #l, which has a higher importance in its own agent #i, will increase increases. Therefore, in the request parameter d i When the value is 1, the task #l selected for the agent #i is... i * The likelihood of other intelligent agents #j being selected in the processing of the control device 100 increases. That is, other intelligent agents #j come to support task #l. i * The possibility of this increases. Therefore, in task #1 i * If not performed, by making the request parameter d i Becoming 1 increases the ability to coordinate task execution through one's own agent #i and other agents #j. i * The possibility of achieving task #l. i * The likelihood of achieving the target increases.

[0161] Furthermore, in response parameter σ iWhen the value is 0, in the processing of the control device 100 of its own intelligent agent #i, the second term of the upper equation on the right side of equation (23) and the lower equation on the right side become 0. Therefore, in the response parameter σ i When the value is 0, in the processing of the control device 100 of its own agent #i, it is difficult to increase the importance of task #l, which has a higher importance among other agents #j (requesting agents). Here, task #l, which has a higher importance among other agents #j, may include task #l selected for other agents #j. j * Based on the above, in response parameter σ i When the value is 0, in the processing of the control device 100 of its own intelligent agent #i, the task #l that is selected for other intelligent agents #j is chosen. j * The likelihood of this happening decreases. Therefore, the agent #i should support the task. j * The likelihood of that decreases.

[0162] Furthermore, in this case, according to the first term of the above equation on the right-hand side of equation (23), the importance φ of the nearby task #l i l Gradually approaching the target value c i l Here, the nearby task #l is the task #l that was selected in the previous control cycle of agent #i and is currently being executed by agent #i. i * As long as the response parameter σ i If the value is 0, the likelihood of continuing to select the nearby task #l increases.

[0163] In contrast, in the response parameter σ i When the value is 1, in the processing of the control device 100 of its own intelligent agent #i, the second term of the upper equation on the right side of equation (23) and the lower equation on the right side are not 0. That is, in the response parameter σ i When the value is 1, in the processing of the control device 100 of its own agent #i, the probability that the importance of task #l, which has a higher importance in the requesting agent #j, will increase increases. Based on the above, in the response parameter σ i When the value is 1, the task #l selected for the requesting agent #j is... j * The likelihood of being selected during the processing of its own intelligent agent #i's control device 100 increases. That is, its own intelligent agent #i goes to support task #l j * The probability increases. Therefore, in tasks #l selected by other requesting agents #j...j * Without performing this operation, by making the response parameter σ i Becoming 1 improves the ability to coordinate task execution through its own agent #i and other requesting agents #j. j * The possibility of achieving task #l. j * The likelihood of achieving the target increases.

[0164] Figure 4 Here is a flowchart illustrating the control method executed by the control device 100 according to Embodiment 1. As described above, the observation information acquisition unit 110 calculates the distances between the surrounding agents 10 and the cargo (task 50) and its own agent 10 (step S102). As described above, the observation information acquisition unit 110 acquires the observation information of its own agent 10 (agent #i). i (Step S110).

[0165] As described above, the action output unit 120 uses strategy π. NN,i This outputs observation information. i Corresponding action a i (Step S120). The request response processing unit 130 performs request response processing (step S130). Specifically, as described above, the request response processing unit 130 processes the request parameter d related to its own agent #i. i and response parameter σ i Perform the calculation.

[0166] The importance processing unit 140 updates the importance of the goods as task 50 (step S140). Specifically, as described above, the importance processing unit 140 updates (calculates) the importance of surrounding tasks #1 (goods) related to its own agent #i. Furthermore, the importance processing unit 140 performs importance processing for goods that have reached their destination (step S142). Specifically, as described above, the importance processing unit 140 sets the importance of task #1 (goods) that has reached its destination and achieved its goal to 0.

[0167] Furthermore, as described above, the task selection unit 150 selects the goods that are most important to its own agent 10 (step S150). As described above, the task execution unit 160 performs processing by having its own agent 10 transport the selected goods (step S160).

[0168] The control device 100 determines whether the distance between the position of all goods and their target position is less than a fixed value (step S170). Here, "the distance between the position of the goods and their target position is less than a fixed value" means that the goods can be considered to have reached the target position. Therefore, the control device 100 determines whether all goods have reached the target position. In addition, the "fixed value" corresponds to δ in equation (25) (for example, δ = 0.05). If the distance between the position of all goods and their target position is less than the fixed value (yes in S170), the process ends. On the other hand, if the distance between the position of all goods and their target position is not less than the fixed value (no in S170), the process returns to S102. Then, the process of S102 to S170 is repeated. This repeated process is performed for each of the above control cycles.

[0169] As described above, the control device 100 according to Embodiment 1 performs processing for calculating request parameters and response parameters based on observation information, and for calculating the importance of each task related to the agent based on the request parameters and response parameters. Then, the control device 100 according to Embodiment 1 selects the task that the agent should perform according to the importance, and controls the agent to perform the selected task. The control device 100 according to Embodiment 1 can be configured in this way to appropriately select the task that the agent should perform based on the importance calculated based on observation information, request parameters, and response parameters. Therefore, even in environments where the task is unknown, it is possible to suppress the situation where a large number of agents are concentrated on a single task, and to achieve actions such as agents supporting tasks that have not yet been performed. Thus, the task is performed reliably. Therefore, even in environments where the task is unknown, the goal of the task can be effectively achieved. Therefore, as a whole, the execution time of the task (total execution time) can be reduced.

[0170] Furthermore, the control device 100 according to Embodiment 1 is configured to calculate request parameters and response parameters based on a strategy learned for each agent. Preferably, the control device 100 according to Embodiment 1 is configured to calculate request parameters and response parameters separately based on the request level and response level, which are input into and output from the strategy using observation information. This allows for the calculation of importance for each agent in a manner that the importance of the task requiring support may increase. Therefore, it is possible to appropriately select the task to be performed for each agent.

[0171] Furthermore, the control device 100 according to Embodiment 1 calculates request parameters indicating a request for support when the request level exceeds a pre-defined threshold and the task being performed or intended to be performed by the agent is not being performed. With this structure, it is possible to appropriately calculate request parameters indicating a request for support when support should be requested for the task being performed or intended to be performed by the agent. Therefore, it is possible to suppress the occurrence of useless requests.

[0172] Furthermore, the control device 100 according to Embodiment 1 calculates response parameters indicating the response request situation when the response level exceeds a pre-defined threshold and the task being performed or intended to be performed by the agent is not being performed. With this structure, it is possible to continue performing the task even when the agent is performing or intending to perform the task. Therefore, it is possible to suppress situations where useless responses are implemented.

[0173] Furthermore, the control device 100 according to Embodiment 1 calculates the importance of each task related to each agent based on the strategy learned for each agent. With this structure, the importance of each task can be appropriately calculated for each agent.

[0174] Furthermore, the control device 100 according to Embodiment 1 calculates the importance of the task corresponding to the observation information associated with the agent based on a target value of the importance of the task corresponding to the observation information, which is input into the strategy and output from the strategy. With this structure, the importance of the task corresponding to the observation information can be calculated for each agent in a manner close to the target value. Therefore, the importance of the task can be calculated appropriately.

[0175] (Implementation Method 2)

[0176] Next, Embodiment 2 will be described. Furthermore, regarding the structure of the control system 1 involved in Embodiment 2, since it is related to... Figure 1 The structure of the control system 1 shown in Embodiment 1 is essentially the same, therefore description is omitted. Furthermore, regarding the structure of the control device 100 in Embodiment 2, since it is similar to... Figure 2 The control device 100 involved in Embodiment 1 is essentially the same in structure, so its description is omitted. In Embodiment 2, task 50 differs from that in Embodiment 1.

[0177] In Implementation 2, Task 50 involves a location containing a large quantity of goods that must be moved. Furthermore, a destination (target location) is set for each type of goods as its moving destination. The objective of Task 50 in Implementation 2 is that all goods present in the location reach their respective destinations. Unlike Implementation 1, in Implementation 2, the goods being moved may be small enough to be moved by a single agent 10. Additionally, the more agents 10 that move the goods in the location (Task 50), the higher the probability that Task 50 will achieve its objective (the probability that all goods in the location will reach their destinations).

[0178] The task 50 involved in Embodiment 2 can be, for example, various rooms within a hospital. Furthermore, it can be configured such that each room, serving as task 50, contains medical records, medications, samples, etc., which are goods to be moved. Additionally, the task 50 involved in Embodiment 2 can also be, for example, a location for storing relief supplies during a disaster. Moreover, relief supplies can also be goods that need to be moved.

[0179] In Embodiment 2, the observation information acquisition unit 110 acquires the observation information as shown by equation (5) in the same manner as in Embodiment 1. i At this time, the observation information acquisition unit 110 calculates the distance between its own agent #i and the location that is the task #l. Specifically, similar to Embodiment 1, the observation information acquisition unit 110 acquires the location that is the task #l. Based on the acquired location of the task #l (location #l), the observation information acquisition unit 110 calculates the distance between its own agent #i and each task #l (location #l). Then, based on the distance between its own agent #i and each task #l (location #l), the observation information acquisition unit 110 determines a predetermined number of nearby tasks that are closer to its own agent #i, and sets information related to the nearby tasks as part of the observation information. In the examples of equations (2) to (5) described above, the nearby tasks are the location #l1 closest to its own agent #i and the second closest location #l2.

[0180] Furthermore, in Implementation 2, o_(l1)^task in Equation (5) can also represent the state of each cargo in location #l1, which is a nearby task, and the destination of each cargo. The state of each cargo can also be the position and speed of each cargo. In addition, o_(l1)^task can also represent the average state (position and speed) of each cargo in location #l1, which is a nearby task, and the average position of the destination of each cargo. The average position can also be the center of gravity (geometric center) of each position (destination). Alternatively, o_(l1)^task can also represent the number of cargo in location #l1, which is a nearby task, and the target number of cargo (i.e., 0). The same applies to o_(l2)^task.

[0181] In Embodiment 2, the policy storage unit 112 stores the policy π (learned model) learned through reinforcement learning in the same manner as in Embodiment 1. The policy π is learned for each agent 10. The policy π for agent #i... NN,i The observation information described above i Let it be the input, and let the action a shown in equation (12) above be... i Set as the output. Furthermore, similar to implementation 1, strategy π... NN,i So that the reward r shown by equation (14) above is obtained. i (t) is used to maximize learning.

[0182] Here, in implementation 2, P is related to the completion degree of task #1. l (t) can also represent the completion rate of the transport of goods at location #1 at time t. The completion rate can, for example, correspond to the ratio of the number of goods transported to the destination to the number of goods initially in location #1. Furthermore, in embodiment 2, Q, related to the progress of task #1... l (t) can also represent the decrease in goods at time t. The decrease in goods can also correspond to, for example, the amount of (moved) goods subtracted from location #l per unit time.

[0183] Figure 5 Here is a flowchart illustrating the control method executed by the control device 100 according to Embodiment 2. As described above, the observation information acquisition unit 110 calculates the distances between the surrounding agents 10 and the location (task 50) and its own agent 10 (step S202). As described above, the observation information acquisition unit 110 acquires the observation information of its own agent 10 (agent #i). i (Step S210). The action output unit 120 uses strategy π in the same manner as in embodiment 1.NN,i Output and observation information o i Corresponding action a i (Step S220).

[0184] The request response processing unit 130 performs request response processing in the same manner as in Embodiment 1 (step S230). Specifically, the request response processing unit 130 may also use the above-described formula (18) to process the request parameter d related to its own agent #i. i The calculation is performed. Furthermore, the request response processing unit 130 can also use the above equation (19) to process the response parameter σ related to its own agent #i. i Perform the calculation.

[0185] The importance processing unit 140 updates the importance of the location serving as task 50 in the same manner as in Embodiment 1 (step S240). That is, the importance processing unit 140 obtains the request parameter d related to agent #j from the control devices 100 of other nearby agents #j in the same manner as in Embodiment 1. j and the importance φ of each task #l (location #l) j l Then, the importance processing unit 140 assigns importance φ to the surrounding tasks #l (locations #l) related to its own agent #i. i l The importance processing unit 140 can also use the above-mentioned equations (23) and (24) to update (calculate) the importance φ of each task #l (location #l) related to its own agent #i. i l Update.

[0186] The importance processing unit 140 performs importance processing on all locations where goods have arrived at their destinations (step S242). Specifically, similar to Embodiment 1, the importance processing unit 140 assigns importance φ to locations #l where all goods have arrived at their destinations and which are related to its own agent #i. i l Set it to 0.

[0187] The task selection unit 150 selects the location with the highest importance to its own agent 10 in the same manner as in Embodiment 1 (step S250). That is, the task selection unit 150 can also use the above-described formula (26) to select the location with the highest importance φ to its own agent #i in the same manner as in Embodiment 1. i l The largest venue #l.

[0188] The task execution unit 160 controls the movement of its own intelligent agent 10 to the selected location and performs cargo handling (step S260). Specifically, the task execution unit 160 moves to the selected location 1... i * The agent #i is controlled in this manner. Furthermore, the task execution unit 160 controls the agent #i when it moves to location l. i * The time will exist in the place l i * The method of transporting goods to the destination is controlled. The method of transporting goods is the same as that described in Implementation Method 1 above.

[0189] The control device 100 determines whether the distance between the location of all goods and their destination in all locations is less than a fixed value (step S270). Here, "the distance between the location of all goods and their destination is less than a fixed value" means that the goods can be considered to have reached their destination. Therefore, the control device 100 determines whether all goods in all locations have reached their destination. In addition, the "fixed value" may also correspond to δ in equation (25). If the distance between the location of all goods and their destination in all locations is less than the fixed value (yes in S270), the processing flow ends. On the other hand, if the distance between the location of all goods and their destination in all locations is not less than the fixed value (no in S270), the processing flow returns to S202. Then, the processing of S202 to S270 is repeated. This repeated processing is performed for each control cycle described above.

[0190] Similar to Embodiment 1, the control device 100 in Embodiment 2 calculates request and response parameters based on observation information and performs processing to calculate the importance of each task related to the agent based on the request and response parameters. Then, the control device 100 in Embodiment 2 selects the task that the agent should perform based on the importance and controls the agent to perform the selected task. Therefore, similar to Embodiment 1, the control device 100 in Embodiment 2 can appropriately select the task that the agent should perform based on the importance calculated based on observation information, request parameters, and response parameters. Thus, even in environments where the task is unknown, it is possible to suppress the situation where a large number of agents are concentrated on a single task and to enable actions such as agents supporting tasks that have not yet been performed. Therefore, the task will be performed reliably. Therefore, even in environments where the task is unknown, the task objective can be effectively achieved. Therefore, overall, the task execution time (total execution time) can be reduced.

[0191] (Examples of modifications to Implementation Method 1 and Implementation Method 2)

[0192] Furthermore, although in Embodiments 1 and 2, the intelligent agent 10 is assumed to be a robot or other machine, the intelligent agent 10 may not be a machine. The intelligent agent 10 may also include both machines and humans. That is, multiple goods can be moved by coordinating between robots and humans. In this case, the human may also carry a communication terminal capable of communicating with the control device 100 associated with the intelligent agent 10. In addition, the intelligent agent 10, as a machine, can be controlled through a process substantially the same as that described in Embodiments 1 and 2 above.

[0193] At this time, the control device 100, acting as the machine agent 10, can also send its own request parameters and the importance of each task to the communication terminal carried by the human. The human can also respond to the agent 10 that sent request parameters indicating a need for support, and support the task being performed by that agent 10, based on the request parameters obtained from other agents 10 and the importance of each task, using their own judgment. Furthermore, the human can use the sense of accomplishment from completing the transport of goods as a guiding principle to independently judge which goods should be transported. Additionally, the human may not request support from other agents 10. That is, the human may also choose not to send request parameters indicating a need for support to other agents 10. This is because, since the timing of human requests cannot be simulated in the learning of strategies related to the robot agent 10, if the human requests support, undesirable results may occur in the actions of the robot agent 10.

[0194] (Implementation Method 3)

[0195] Next, Embodiment 3 will be described. Furthermore, regarding the structure of the control system 1 involved in Embodiment 3, since it is related to... Figure 1 The structure of the control system 1 shown in Embodiment 1 is essentially the same, therefore description is omitted. Furthermore, regarding the structure of the control device 100 in Embodiment 3, since it is similar to... Figure 2The structure of the control device 100 involved in Embodiment 1 is essentially the same, so its description is omitted. In Embodiment 3, task 50 differs from the embodiments described above. Here, in the embodiments described above, the intelligent agent 10 moves the goods to achieve the goal of task 50. In contrast, in Embodiment 3, the intelligent agent 10 may not move the goods when performing task 50. Specific examples of task 50 in Embodiment 3 will be described later. Similar to Embodiment 1, the intelligent agent 10 autonomously performs actions within the environment through control implemented by the control device 100. Furthermore, in Embodiment 3, the monitoring device 60 may not monitor task 50. Each intelligent agent 10 may also monitor (detect) the state of task 50.

[0196] Figure 6 Here is a flowchart illustrating the control method executed by the control device 100 according to Embodiment 3. The observation information acquisition unit 110 calculates the distances between the surrounding agents 10 and the task 50 and its own agent 10 in the same manner as in S102 and S202 (step S302). The observation information acquisition unit 110 acquires the observation information of its own agent 10 (agent #i) in the same manner as in S110 and S210. i (Step S310). Regarding observation information o i This will be described later. The action output unit 120 uses strategy π in the same manner as S120 and S220. NN,i and output observation information. i Corresponding action a i (Step S320). Regarding strategy π NN,i Related compensation r i (t), which will be discussed later.

[0197] The request response processing unit 130 performs request response processing in the same manner as S130 and S230 (step S330). Specifically, the request response processing unit 130 may also use the above formula (18) to process the request parameter d related to its own agent #i. i The calculation is performed. Furthermore, the request response processing unit 130 can also use the above equation (19) to process the response parameter σ related to its own agent #i. i Perform the calculation.

[0198] The importance processing unit 140 updates the importance of task 50 in the same manner as in S140 and S240 (step S340). Specifically, the importance processing unit 140 obtains the request parameter d related to agent #j from the control device 100 of other nearby agents #j in the same manner as in the above-described embodiment.j and the importance φ of each task #l j l Then, the importance processing unit 140 assigns importance φ to the surrounding tasks #l related to its own agent #i in the same manner as in the embodiment described above. i l The importance processing unit 140 can also use the above equations (23) and (24) to update (calculate) the importance φ of each task #l related to its own agent #i. i l Update.

[0199] The importance processing unit 140 processes the importance of completed tasks in the same manner as in S142 and S242 (step S342). Specifically, similar to the implementation described above, the importance processing unit 140 assigns importance φ to the task #l that has achieved its objective and is related to its own agent #i. i l Set it to 0.

[0200] The task selection unit 150 selects the task 50 with the highest importance to its own agent 10 in the same manner as S150 and S250 (step S350). Specifically, the task selection unit 150 can also use the above-described formula (26) to select the task with the highest importance φ to its own agent #i in the same manner as the implementation described above. i l The biggest task #l.

[0201] Similar to S160 and S260, the task execution unit 160 controls the selected task 50 by having its own agent 10 execute the selected task 50 (step S360). Specifically, the task execution unit 160 moves to the selected task 50. i * The agent #i is controlled by its position. Furthermore, the task execution unit 160 controls the agent #i when it moves to task l. i * When the location is to execute the task l i * This will be controlled in a certain way. A specific example of Task 50 will be described later.

[0202] The control device 100 determines whether all tasks 50 have been completed (step S370). If all tasks 50 have been completed (Yes in S370), the processing flow ends. On the other hand, if all tasks 50 have not been completed (No in S370), the processing flow returns to S302. Then, the processing of S302 to S370 is repeated. This repetition of processing is performed for each control cycle described above.

[0203] Similar to Embodiment 1, the control device 100 in Embodiment 3 performs processing for calculating request parameters and response parameters based on observation information, and for calculating the importance of each task related to the agent based on the request parameters and response parameters. Then, the control device 100 in Embodiment 3 selects the task that the agent should perform based on the importance, and controls the agent to perform the selected task. Therefore, similar to Embodiment 1, the control device 100 in Embodiment 3 can appropriately select the task that the agent should perform based on the importance calculated based on observation information, request parameters, and response parameters. Thus, even in environments where the task is unknown, unnecessary concentration, such as multiple agents focusing on a single task, can be suppressed, and actions such as agents supporting tasks that have not yet been performed can be implemented. Therefore, the task can be performed reliably. Therefore, even in environments where the task is unknown, the task objective can be effectively achieved. Therefore, overall, the task execution time (total execution time) can be reduced.

[0204] <Specific Example 1 of Implementation Method 3>

[0205] In Specific Example 1, the method described in this embodiment is applied to maintenance. In Specific Example 1, maintenance (repair) of a structure is performed by multiple intelligent agents 10, which are machines such as robots. Furthermore, in Specific Example 1, the multiple intelligent agents 10 perform a wide-ranging inspection of the structure. Here, in Specific Example 1, task 50 is to inspect each part. Furthermore, in Specific Example 1, the goal of task 50 is to perform a multi-faceted inspection of each inspection part. In addition, the multiple intelligent agents 10 can each have different functions. Therefore, the multiple intelligent agents 10 can be composed of different types of intelligent agents 10. By using multiple intelligent agents 10 of different types, a multi-faceted inspection can be achieved. Therefore, the more intelligent agents 10 there are, the higher the possibility of achieving the goal of task 50.

[0206] For example, one intelligent agent 10 may be a search robot that searches for abnormal areas. Furthermore, other intelligent agents 10 may be processing robots that handle abnormalities. Additionally, one intelligent agent 10 (the search robot) may have a camera that captures images of the area under inspection and determines the severity of the abnormality based on the captured images. Furthermore, one intelligent agent 10 may have the capability to perform a first non-destructive inspection (e.g., infrared survey). Furthermore, other intelligent agents 10 may have the capability to perform a second non-destructive inspection (e.g., ultrasonic testing). Furthermore, other intelligent agents 10 may have the capability to perform a third non-destructive inspection (e.g., X-ray transmission testing). Furthermore, other intelligent agents 10 may have the capability to perform a fourth non-destructive inspection (e.g., eddy current testing).

[0207] In specific example 1, the observation information acquisition unit 110 acquires observation information as shown in equation (5) in the same manner as the implementation described above. i (S310). At this time, the observation information acquisition unit 110 calculates the distance between its own agent #i and the task #l (maintenance part #l) (S302). Specifically, similar to Embodiment 1, the observation information acquisition unit 110 acquires the location of the maintenance part that is the task #l. The observation information acquisition unit 110 calculates the distance between its own agent #i and each task #l (maintenance part #l) based on the acquired location of the task #l (maintenance part #l). Then, the observation information acquisition unit 110 determines a predetermined number of nearby tasks that are closer to its own agent #i based on the distance between its own agent #i and each task #l (maintenance part #l), and sets information related to the nearby tasks as part of the observation information. In the examples of equations (2) to (5) described above, the nearby tasks are the maintenance part #l1 that is closest to its own agent #i and the second closest maintenance part #l2.

[0208] Furthermore, in Specific Example 1, o_(l1)^task in Equation (5) can also represent the state of the maintenance location #l1, which is a nearby task, and the end condition of the maintenance. The state of each maintenance location can also be the location of the maintenance location, the content of the maintenance performed, and the severity of the anomaly. The end condition of the maintenance can also be that all types of agents 10 have arrived at the maintenance location and performed all kinds of checks (multi-faceted checks). The same applies to o_(l2)^task.

[0209] In Specific Example 1, the policy storage unit 112 stores the policy π (learned model) learned through reinforcement learning in the same manner as in the implementation described above. The policy π is learned for each agent 10. The policy π for agent #i... NN,i The observation information described above i Let it be the input, and let the action a shown in equation (12) above be... i Set as the output. Furthermore, similar to implementation 1, strategy π... NN,i So that the reward r shown by equation (14) above is obtained. i (t) is used to maximize learning.

[0210] Here, in specific example 1, P is related to the completion degree of task #1. l (t) can also represent the completion rate of maintenance at maintenance location #l at time t. The completion rate can, for example, correspond to the number of agent types 10 that have reached the critical maintenance location #l and performed processing. Alternatively, the completion rate can also be the number of maintenance items performed. Furthermore, in Specific Example 1, Q is related to the progress of task #l. l (t) can also be the maintenance progress of maintenance part #l at time t. The maintenance progress can also correspond to the number of agents 10 reaching the critical maintenance part #l per unit time. Alternatively, the inspection progress can also be the number of maintenance items performed per unit time.

[0211] Furthermore, in Specific Example 1, the Request Response Processing Unit 130, in the same manner as the implementation described above, responds according to strategy π. NN,i Output action a i Thus, the request parameter d related to agent #i is... i and response parameter σ i Calculations are performed (S330). Then, the request response processing unit 130 sends the request parameter d to the control device 100 of the surrounding intelligent agent 10. i Furthermore, in Specific Example 1, the importance processing unit 140 assigns importance φ to peripheral tasks #l (maintenance parts #l) related to its own agent #i in the same manner as in the implementation described above. i l The importance processing unit 140 can also use the above-mentioned equations (23) and (24) to update (calculate) the importance φ of each task #l (maintenance part #l) related to its own agent #i. i lUpdate. Furthermore, in Specific Example 1, the task selection unit 150, in the same manner as the implementation described above, uses the above formula (26) to select all tasks #l with importance φ. i l The most important task #l (maintenance part #l) is the task that the agent #i should perform. i * (S350).

[0212] In this specific example 1, the request response processing unit 130 can also send request parameters to the control device 100 of an agent 10 of a different type than its own agent #i. This increases the likelihood that an agent 10 of a different type than its own agent #i will arrive at the maintenance location. Conversely, it decreases the likelihood that an agent 10 of the same type as its own agent #i will arrive at the maintenance location. That is, if agent #i, acting as a search robot, detects a serious maintenance location #l, it is envisioned that the request level output from the strategy in the control device 100 of that agent #i will increase. Then, the control device 100 of agent #i, acting as a search robot, will send request parameters to the control device 100 of different types of agents 10 (such as agents 10 performing non-destructive inspections). i The request parameter is 1. Therefore, it is conceivable that the importance of the repair location #1 increases in control devices 100 for different types of agents 10 (such as agents 10 performing non-destructive inspections). Consequently, in control devices 100 for different types of agents 10 (such as agents 10 performing non-destructive inspections), the probability that the repair location #1 is selected and that different types of agents 10 reach the repair location #1 increases. Therefore, the probability of achieving the objective of task 50 increases. Furthermore, if information about which repair location's inspection has not been completed is added to the observation information, the agent that has obtained the observation information can determine whether it should actively respond to the request for support to that repair location.

[0213] Furthermore, in Specific Example 1, the task execution unit 160 controls the process by having its own agent #i execute the selected task #l (maintenance part #l) (S360). Specifically, the task execution unit 160 moves the agent #i to task #l. i * (Inspection section #l) i * The task execution unit 160 controls its own agent #i in a manner that conforms to the function of its agent #i. Then, the control device 100 of each agent 10 performs processing to perform multi-faceted inspections on all inspection parts.

[0214] <Specific Example 2 of Implementation Method 3>

[0215] In Specific Example 2, the method described in this embodiment is applied to monitoring, patrolling, and security. In Specific Example 2, environmental security is implemented using multiple intelligent agents 10, which are machines such as robots. Specifically, the multiple intelligent agents 10 patrol the environment and perform processes related to monitoring, patrolling, and security. The multiple intelligent agents 10 patrol the environment and search for issues existing in the environment. Furthermore, in Specific Example 2, when an issue is found, the control device 100 of the intelligent agent 10 that found the issue sends information related to the found issue to the control devices 100 of the surrounding intelligent agents 10. Here, in Specific Example 2, task 50 is "the issue that was searched." Furthermore, in Specific Example 2, the goal of task 50 is to resolve the situation where the searched issue is found. The more intelligent agents 10 there are, the higher the probability of achieving the goal of task 50.

[0216] Furthermore, in Specific Example 2, multiple agents 10 may also have the function of processing the searched issue. Additionally, as in Specific Example 1, multiple agents 10 may have different functions. For example, if the searched issue is "removal of bulky waste," an agent 10 capable of transporting bulky waste can remove it. Furthermore, if the searched issue is "protection of criminals," an agent 10 capable of protecting criminals can protect them. Furthermore, if the searched issue is "handling of lost people," an agent 10 capable of providing road guidance can handle the situation of lost people. In the following description, "searched issue" may sometimes be referred to simply as "issue."

[0217] In specific example 2, the observation information acquisition unit 110 acquires observation information as shown in equation (5) in the same manner as the implementation described above. i(S310). At this time, the observation information acquisition unit 110 calculates the distance between its own agent #i and the task #l (topic #l) (S302). Specifically, similar to Embodiment 1, the observation information acquisition unit 110 acquires the position of the "searched topic" which is the task #l. The position of the "searched topic" can also be the position of the agent 10 that searched for the searched topic at the time of the search. The observation information acquisition unit 110 calculates the distance between its own agent #i and each task #l (topic #l) based on the acquired position of the task #l (topic #l). Then, the observation information acquisition unit 110 determines a predetermined number of nearby tasks that are closer to its own agent #i based on the distance between its own agent #i and each task #l (topic #l), and sets information related to the nearby tasks as part of the observation information. In the examples of equations (2) to (5) described above, the nearby tasks are the topic #l1 that is closest to its own agent #i and the second closest topic #l2.

[0218] Furthermore, in specific example 2, o_(l1)^task in equation (5) can also represent the state of task #l1 as a nearby task and the end condition of the task. The state of each task can also be the position of the task, the degree of resolution of the task, the quality of the task, and the type of the task. The end condition of the task can also be the status of the task resolution. The same applies to o_(l2)^task.

[0219] In specific example 2, the policy storage unit 112 stores the policy π (learned model) learned through reinforcement learning in the same manner as in the implementation described above. The policy π is learned for each agent 10. The policy π for agent #i... NN,i The observation information described above i Let it be the input, and let the action a shown in equation (12) above be... i Set as the output. Furthermore, similar to implementation 1, strategy π... NN,i So that the reward r shown by equation (14) above is obtained. i (t) is used to maximize learning.

[0220] Here, in specific example 2, P is related to the completion degree of task #1. l (t) can also represent the degree of completion of task #l at time t. The degree of completion can, for example, correspond to the situation where the agent 10 capable of handling the task finishes processing. Furthermore, in specific example 2, Q related to the progress of task #l... l(t) can also be the progress of processing the task #l at time t. The progress of task processing can also correspond to the situation where the agent 10 capable of processing the task is performing processing.

[0221] Furthermore, in Specific Example 2, the Request Response Processing Unit 130, in the same manner as the implementation described above, responds according to strategy π. NN,i Output action a i To request the parameter d related to agent #i i and response parameter σ i Calculations are performed (S330). Then, the request response processing unit 130 sends the request parameter d to the control device 100 of the surrounding intelligent agent 10. i At this time, as in variations of Embodiment 1 and Embodiment 2, the control device 100 can also send request parameters to a terminal carried by a human. Thus, a human can also process the problem. Furthermore, as in Specific Example 1, the control device 100 can also send request parameters to the control device 100 of an intelligent agent 10 of a different type than its own intelligent agent #i.

[0222] Furthermore, in specific example 2, the importance processing unit 140 assigns importance φ to the surrounding tasks #l (topics #l) related to its own agent #i in the same manner as in the implementation described above. i l The importance processing unit 140 can also use the above-mentioned equations (23) and (24) to update (calculate) the importance φ of each task (topic) related to its own agent #i. i l Update. Furthermore, in specific example 2, the task selection unit 150, in the same manner as the implementation described above, uses the above formula (26) to determine the importance φ of all tasks #l. i l The biggest task (problem) is choosing the task that the agent #i should perform. i * (S350).

[0223] Furthermore, a scenario is envisioned where, if agent #i, which has identified issue #l, lacks the capability to process that issue, the request level from the policy output increases within the control device 100 of agent #i. Then, the control device 100 of agent #i sends a request to the control devices 100 of neighboring agents 10. iThe request parameter is set to 1. Then, it can be envisioned that, within the control device 100 of the surrounding agents 10 capable of processing the task, by acquiring observation information indicating the quality and type of the task, as well as the aforementioned request parameter, the importance of task #1 is increased. Therefore, within the control device 100 of the agents 10 capable of processing the task, the probability that task #1 has been selected and that the agent 10 capable of processing the task has reached the location of task #1 increases. Furthermore, if information about which task's processing has not yet ended is added to the observation information, it is possible to determine, on the side of the agent that has acquired the observation information, whether a proactive response to the request for support for that task should be initiated.

[0224] Furthermore, in specific example 2, the task execution unit 160 controls the process by causing its own agent #i to execute the selected task #l (task #l) (S360). Specifically, the task execution unit 160 moves the agent #i to task #l. i * (Topic #l) i * The task execution unit 160 controls its own agent #i to perform the processing of the tasks it can handle. Then, the control devices 100 of each agent 10 implement processing to solve all the searched tasks.

[0225] <Specific Example 3 of Implementation Method 3>

[0226] In Specific Example 3, the method described in this embodiment is applied to coexistence with nature. In Specific Example 3, multiple intelligent agents 10, acting as robots or other machines, are used to monitor and control animal activities. This reduces agricultural losses while preventing animals from encroaching on the farm, thus achieving a sustainable ecosystem.

[0227] In Specific Example 3, multiple agents 10 detect animals by detecting objects moving on or around the farm. Furthermore, each agent 10 can have different functions. That is, as in Specific Example 1, the multiple agents 10 can be composed of agents of different types. In this case, one agent 10 can also have the function of searching for animals. Additionally, other agents 10 can also have the function of causing animals to leave the farm. Alternatively, each agent 10 can have both the function of searching for animals and the function of causing animals to leave the farm. That is, the multiple agents 10 can also be agents of the same type. Furthermore, the control device 100 of the agent 10 that has detected an animal sends animal-related information (animal location, etc.) to the control devices 100 of the other agents 10. Here, in Specific Example 3, task 50 is the detection of each animal. Furthermore, in Specific Example 3, the goal of task 50 is the departure of the animals. The more agents 10 there are, the higher the probability of achieving the goal of task 50.

[0228] In specific example 3, the observation information acquisition unit 110 acquires observation information as shown in equation (5) in the same manner as the implementation described above. i (S310). At this time, the observation information acquisition unit 110 calculates the distance between its own agent #i and the task #l (animal #l) (S302). Specifically, similar to Embodiment 1, the observation information acquisition unit 110 acquires the position of the animal that is the task #l. The observation information acquisition unit 110 calculates the distance between its own agent #i and each task #l (animal #l) based on the acquired position of the task #l (animal #l). Then, the observation information acquisition unit 110 determines a predetermined number of nearby tasks that are closer to its own agent #i based on the distance between its own agent #i and each task #l (animal #l), and sets information related to the nearby tasks as part of the observation information. In the examples of equations (2) to (5) described above, the nearby tasks are the animal #l1 that is closest to its own agent #i and the second closest animal #l2.

[0229] Furthermore, in specific example 3, o_(l1)^task in equation (5) can also represent the state of animal #l1 as a nearby task and the animal #l1 leaving its destination (end point). The animal's state can also be the animal's position and speed. The animal #l1 leaving its destination can also correspond to the animal's original territory. The same applies to o_(l2)^task.

[0230] In specific example 3, the policy storage unit 112 stores the policy π (learned model) learned through reinforcement learning in the same manner as in the implementation described above. The policy π is learned for each agent 10. The policy π for agent #i... NN,i The observation information described above i Let it be the input, and let the action a shown in equation (12) above be... i Set as the output. Furthermore, similar to implementation 1, strategy π... NN,i So that the reward r shown by equation (14) above is obtained. i (t) is used to maximize learning.

[0231] Here, in specific example 3, P is related to the completion degree of task #1. l (t) can also be set as the distance at time t up to the boundary of the area to be protected from animals (the area where animals are not allowed to invade). Furthermore, in specific example 3, Q is related to the progress of task #1. l (t) can also be the speed at which animal #l moves toward the destination at time t.

[0232] Furthermore, in specific example 3, the request response processing unit 130, in the same manner as the implementation described above, responds according to strategy π. NN,i Output action a i Thus, the request parameter d related to agent #i is... i and response parameter σ i Calculations are performed (S330). Then, the request response processing unit 130 sends the request parameter d to the control device 100 of the surrounding intelligent agent 10. i At this time, as in variations of Embodiment 1 and Embodiment 2, the control device 100 can also send request parameters to a terminal carried by a human. Thus, the human can also make the animal leave.

[0233] Furthermore, in specific example 3, the importance processing unit 140 assigns importance φ to the surrounding tasks #l (animals #l) related to its own agent #i in the same manner as in the implementation described above. i l The importance processing unit 140 can also use the above-mentioned equations (23) and (24) to update (calculate) the importance φ of each task #l (animal #l) related to its own agent #i. i l Update. Furthermore, in specific example 3, the task selection unit 150, in the same manner as the implementation described above, uses the above formula (26) to determine the importance φ of all tasks #l. il The biggest task #l (animal #l) is to choose the task that #i, as its intelligent agent, should perform #l i * (S350).

[0234] Furthermore, when multiple agents 10 are composed of agents 10 of different types, similar to Example 1, the request response processing unit 130 can also send request parameters to the control device 100 of an agent #i that is different from its own agent #i. Therefore, the likelihood of an agent 10 of a different type from its own agent #i reaching the animal is increased. Thus, similar to Example 1, when an agent #i with the function of searching for animals detects an animal, the likelihood of an agent 10 with the function of causing the animal to leave the farm reaching the animal is increased.

[0235] Furthermore, in specific example 3, the task execution unit 160 performs control by causing its own agent #i to execute the selected task #l (animal control) (S360). Specifically, the task execution unit 160 moves the agent #i to task #l. i * (Animal #l) i * At the location of the animal. The task execution unit 160 controls its own agent #i by executing actions to make the animal leave (driving the animal away). Then, the control device 100 of each agent 10 performs processing to execute actions to make the animal leave (drive it away) for all animals.

[0236] <Specific Example 4 of Implementation Method 3>

[0237] In Specific Example 4, the method described in this embodiment is applied to the provision of a variety of services. In Specific Example 4, multiple intelligent agents 10, which act as machines such as robots, are used to support people living in the environment. This improves people's comfort. Specifically, the multiple intelligent agents 10 circulate in the environment and perform processing to meet people's needs. The intelligent agents 10 circulate in the environment to solve the problems requested by people. Furthermore, in Specific Example 4, when a problem is requested, the control device 100 of the intelligent agent 10 with the requested problem sends information related to the requested problem to the control devices 100 of the surrounding intelligent agents 10. In Specific Example 4, task 50 is the "requested problem". Furthermore, in Specific Example 4, the goal of task 50 is to solve the requested problem. The more intelligent agents 10 there are, the higher the probability of achieving the goal of task 50.

[0238] Furthermore, in Specific Example 4, multiple intelligent agents 10 may also possess the function of processing the requested task. Additionally, as in Specific Example 1, multiple intelligent agents 10 may also possess different functions. For example, if the requested task is "removal of bulky waste," an intelligent agent 10 capable of handling bulky waste can remove it. Furthermore, if the requested task is "protection of criminals," an intelligent agent 10 capable of protecting criminals can protect them. Furthermore, if the requested task is "handling of lost persons," an intelligent agent 10 capable of providing road guidance can handle the situation of lost persons. In the following description, the "requested task" may sometimes be referred to simply as "task."

[0239] In specific example 4, the observation information acquisition unit 110 acquires observation information as shown in equation (5) in the same manner as the implementation described above. i (S310). At this time, the observation information acquisition unit 110 calculates the distance between its own agent #i and the task #l (issue #l) (S302). Specifically, similar to Embodiment 1, the observation information acquisition unit 110 acquires the position of the "required issue" which is the task #l. The position of the "required issue" can also be the position of the agent 10 that was asked for the issue at the time of the request. The observation information acquisition unit 110 calculates the distance between its own agent #i and each task #l (issue #l) based on the acquired position of the task #l (issue #l). Then, the observation information acquisition unit 110 determines a predetermined number of nearby tasks that are closer to its own agent #i based on the distance between its own agent #i and each task #l (issue #l), and sets information related to the nearby tasks as part of the observation information. In the examples of equations (2) to (5) described above, the nearby tasks are the issue #l1 that is closest to its own agent #i and the second closest issue #l2.

[0240] Furthermore, in specific example 4, o_(l1)^task in equation (5) can also represent the state of task #l1 as a nearby task and the end condition of the task. The state of each task can also be the position of the task, the degree of solution of the task, the quality of the task, and the type of the task. The end condition of the task can also be the status of solving the task. The same applies to o_(l2)^task.

[0241] In specific example 4, the policy storage unit 112 stores the policy π (learned model) learned through reinforcement learning in the same manner as in the implementation described above. The policy π is learned for each agent 10. The policy π for agent #i... NN,i The observation information described above iLet it be the input, and let the action a shown in equation (12) above be... i Set as the output. Furthermore, similar to implementation 1, strategy π... NN,i So that the reward r shown by equation (14) above is obtained. i (t) is used to maximize learning.

[0242] Here, in specific example 4, P is related to the completion degree of task #1. l (t) can also represent the degree of completion of task #l at time t. The degree of completion can, for example, correspond to the situation where agent 10, capable of processing the task, has finished processing. Furthermore, in specific example 4, Q, related to the progress of task #l... l (t) can also represent the progress of processing the task #l at time t. The progress of task processing can also correspond to the situation where the agent 10, which is capable of processing the task, is performing processing.

[0243] Furthermore, in specific example 4, the request response processing unit 130, in the same manner as the implementation described above, responds according to strategy π. NN,i Output action a i Thus, the request parameter d related to agent #i is... i and response parameter σ i Calculations are performed (S330). Then, the request response processing unit 130 sends the request parameter d to the control device 100 of the surrounding intelligent agent 10. i At this time, as in variations of Embodiment 1 and Embodiment 2, the control device 100 can also send request parameters to a terminal carried by a human. Thus, a human can also process the problem. Furthermore, as in Specific Example 1, the control device 100 can also send request parameters to the control device 100 of an intelligent agent 10 of a different type than its own intelligent agent #i.

[0244] Furthermore, in specific example 4, the importance processing unit 140 assigns importance φ to the surrounding tasks #l (topics #l) related to its own agent #i in the same manner as in the implementation described above. i l The importance processing unit 140 can also use the above-mentioned equations (23) and (24) to update (calculate) the importance φ of each task (topic) related to its own agent #i. i l Update. Furthermore, in specific example 4, the task selection unit 150, in the same manner as the implementation described above, uses the above formula (26) to determine the importance φ of all tasks #l. il The biggest task (problem) is choosing the task that the agent #i should perform. i * (S350).

[0245] Furthermore, a scenario is envisioned where, if agent #i, which has been found to handle issue #l, does not possess the capability to process that issue, the request level output from the policy in the control device 100 of agent #i increases. Then, the control device 100 of agent #i sends a request to the control devices 100 of neighboring agents 10. i The request parameter is set to 1. Then, it can be envisioned that, within the control device 100 of the surrounding agents 10 capable of processing the task, by acquiring observation information indicating the quality and type of the task, as well as the aforementioned request parameter, the importance of task #1 is increased. Therefore, within the control device 100 of the agents 10 capable of processing the task, the probability that task #1 has been selected and that the agent 10 capable of processing the task has reached the location of task #1 increases. Furthermore, if information about which task's processing has not yet ended is added to the observation information, it is possible to determine, on the side of the agent that has acquired the observation information, whether a proactive response to the request for support for that task should be initiated.

[0246] Furthermore, in specific example 4, the task execution unit 160 controls the process by causing its own agent #i to execute the selected task #l (task #l) (S360). Specifically, the task execution unit 160 moves the agent #i to task #l. i * (Topic #l) i * The task execution unit 160 controls its own agent #i to perform the processing of the tasks that can be performed. Then, the control device 100 of each agent 10 performs processing to solve all the required tasks.

[0247] <Specific Example 5 of Implementation Method 3>

[0248] In Specific Example 5, the method described in this embodiment is applied to handling an event. In Specific Example 5, multiple intelligent agents 10, which are machines such as robots, control the flow of people during the event. Specifically, the multiple intelligent agents 10 search for the flow of people (crowds) that should be guided, and organize the flow of people (crowds) by guiding them to the positions where they should stay. More specifically, for example, the multiple intelligent agents 10 hold a rope, and by moving each intelligent agent 10 to a predetermined position, the area can be divided by the rope. Thus, the flow of people is organized by guiding them to the divided areas. Furthermore, in Specific Example 5, the control device 100 of the intelligent agent 10 that has searched for the flow of people that should be guided can also send information related to the searched flow of people to the control devices 100 of the surrounding intelligent agents 10.

[0249] Here, in Specific Example 5, Task 50 is "the gathering of people (or simply "people flow")". Furthermore, in Specific Example 5, the objective of Task 50 is to guide the people flow to where they should stay. Additionally, by having a large number of agents 10 organize the people flow, the variations in the divisions increase, and the size of the divided areas increases. Therefore, the more agents 10 there are, the higher the probability of achieving the objective of Task 50.

[0250] In specific example 5, the observation information acquisition unit 110 acquires observation information as shown by equation (5) in the same manner as the implementation described above. i (S310). At this time, the observation information acquisition unit 110 calculates the distance between its own agent #i and task #l (people flow #l) (S302). Specifically, similar to Embodiment 1, the observation information acquisition unit 110 acquires the position of the people flow that is task #l. The observation information acquisition unit 110 calculates the distance between its own agent #i and each task #l (people flow #l) based on the acquired position of task #l (people flow #l). Then, the observation information acquisition unit 110 determines a predetermined number of nearby tasks that are closer to its own agent #i based on the distance between its own agent #i and each task #l (people flow #l), and sets information related to the nearby tasks as part of the observation information. In the example of equations (2) to (5) described above, the nearby tasks are the people flow #l1 that is closest to its own agent #i and the second closest people flow #l2.

[0251] Furthermore, in specific example 5, o_(l1)^task in equation (5) can also represent the state of the flow of people #l1, which is a nearby task, and the destination (the position where the flow should stop). The state of the flow can also be the position of the flow and the speed of the flow. The same applies to o_(l2)^task.

[0252] In specific example 5, the policy storage unit 112 stores the policy π (learned model) learned through reinforcement learning in the same manner as in the implementation described above. The policy π is learned for each agent 10. The policy π for agent #i... NN,i The observation information described above i Let it be the input, and let the action a shown in equation (12) above be... i Set as the output. Furthermore, similar to implementation 1, strategy π... NN,i So that the reward r shown by equation (14) above is obtained. i (t) is used to maximize learning.

[0253] Here, in specific example 5, P is related to the completion degree of task #1. l (t) can also represent whether the flow of people #l has reached its destination (the position where it should stay) at time t. Furthermore, in specific example 5, Q is related to the progress of task #l. l (t) can also be the speed at which the flow of people #l moves toward the destination at time t.

[0254] Furthermore, in Specific Example 5, the Request Response Processing Unit 130, in the same manner as the implementation described above, responds according to strategy π. NN,i Output action a i Thus, the request parameter d related to agent #i is... i and response parameter σ i Calculations are performed (S330). Then, the request response processing unit 130 sends the request parameter d to the control device 100 of the surrounding intelligent agent 10. i At this time, as in the variations of Embodiment 1 and Embodiment 2, the control device 100 can also send request parameters to a terminal carried by a human. Thus, a human can also manage the flow of people.

[0255] Furthermore, in specific example 5, the importance processing unit 140 assigns importance φ to the surrounding tasks #l (people flow #l) related to its own agent #i in the same manner as in the implementation described above. i l The importance processing unit 140 can also use the above-mentioned equations (23) and (24) to update (calculate) the importance φ of each task #l (people flow #l) related to its own agent #i. i l Update. Furthermore, in specific example 3, the task selection unit 150, in the same manner as the implementation described above, uses the above formula (26) to determine the importance φ of all tasks #l.i l The biggest task #l (people flow #l) is the task that the chosen intelligent agent #i should perform. i * (S350).

[0256] Furthermore, in specific example 5, the task execution unit 160 controls the process by having its own agent #i execute the selected task #l (people flow management) (S360). Specifically, the task execution unit 160 moves the agent #i to task #l. i * (human flow) i * The task execution unit 160 controls its own agent #i to perform the organization of the flow of people (guiding people to their proper places). Then, the control devices 100 of each agent 10 perform processing to organize (guide) the flow of people for all of them.

[0257] (Modified example)

[0258] Furthermore, this embodiment is not limited to the above-described embodiment, and appropriate modifications can be made without departing from the main idea. For example, the order of the various steps (processes) in the flowchart described above can be appropriately changed. In addition, one or more steps (processes) in the flowchart described above can be omitted.

[0259] Furthermore, although the agent 10 and task 50 are assumed to exist in physical space in the embodiments described above, this is not a limitation. The agent 10 and task 50 may also exist in a virtual space implemented by simulation, for example.

[0260] The program described above includes, when read by a computer, a set of commands (or software code) used to cause the computer to perform one or more functions described in the embodiments. The program may also be stored on a non-transitory computer-readable medium or a physical storage medium. By way of example, and not limitation, a computer-readable medium or a physical storage medium includes: random-access memory (RAM), read-only memory (ROM), cache memory, solid-state drive (SSD) or other memory technologies, CD-ROM (read-only optical disc), digital versatile disk (DVD), Blu-ray disc or other optical disc storage, magnetic tape, magnetic tape, disk storage, or other magnetic storage devices. The program may also be transmitted on a transient computer-readable medium or a communication medium. By way of example, and not limitation, a transient computer-readable medium or communication medium includes electrical, optical, acoustic, or other forms of transmission signals.

[0261] As will be apparent from the disclosure described herein, embodiments of this disclosure can be varied in many ways. Such variations should not be considered as departing from the spirit and scope of this disclosure, and all such modifications, which will be apparent to those skilled in the art, are intended to be included within the scope of the appended claims.

Claims

1. A control device for controlling an intelligent agent performing a task, wherein, For the task in question, the more agents that perform the task, the higher the probability of achieving the task's objective. Furthermore, for the task in question, there are multiple agents in the environment. The control device has: The request response processing unit calculates request parameters related to whether to request support and response parameters related to whether to respond to requests from other intelligent agents, based on observation information related to the intelligent agent, other intelligent agents in the vicinity of the intelligent agent, and the task. An importance processing unit implements a process for calculating the importance of each of the tasks associated with the agent, based at least on the request parameters of the other agents and the response parameters of the agent. The task selection unit selects the task that the agent should perform based on the importance of the task. A task execution unit controls the manner in which the agent performs the selected task. The importance processing unit calculates the importance of the task corresponding to the observation information associated with that agent based on a target value of the task importance corresponding to the observation information, which is output from the policy learned by each agent by inputting the observation information into the policy. The importance processing unit calculates the importance of the task corresponding to the observation information associated with the agent, in a manner that, if the response parameter indicates that no response is needed from other agents, the importance of the task being performed by that agent is close to the target value. When the response parameter indicates a response to a request from another agent, and the request parameter of the other agent indicates a request for support, the importance processing unit calculates the importance of the task corresponding to the observation information associated with that agent, based on the difference between the importance of the task being performed by the other agent and the importance associated with that agent, in a way that increases the importance of the task being performed by the other agent.

2. The control device as claimed in claim 1, wherein, The request-response processing unit calculates the request parameters and the response parameters based on the strategy.

3. The control device as described in claim 2, wherein, The request-response processing unit calculates the request parameters and the response parameters based on the request level and response level output from the strategy after inputting the observation information into the strategy.

4. The control device as described in claim 3, wherein, The request response processing unit calculates the request parameters indicating a request for support when the request level exceeds a predefined threshold and the task that the agent is performing or wants to perform has not been performed.

5. The control device as described in claim 3, wherein, The request response processing unit calculates the response parameters indicating the status of the response request when the response level exceeds a predefined threshold and the task that the agent is performing or wants to perform has not been performed.

6. A control system, which is a system for distributed control of multiple intelligent agents performing tasks, wherein, For the task in question, the more agents that perform the task, the higher the probability of achieving the task's objective. Furthermore, for the task in question, there are multiple agents in the environment. The control system has multiple control devices as described in claim 1, each capable of controlling multiple intelligent agents. The multiple control devices each have: The request response processing unit calculates request parameters related to whether to request support and response parameters related to whether to respond to requests from other intelligent agents, based on observation information related to the intelligent agent of the control device, other intelligent agents around the intelligent agent, and the task. An importance processing unit implements a process for calculating the importance of each of the tasks associated with the agent, based at least on the request parameters of the other agents and the response parameters of the agent. The task selection unit selects the task that the agent should perform based on the importance of the task. The task execution unit controls the manner in which the agent performs the selected task.

7. A control method for controlling an intelligent agent performing a task, wherein, For the task in question, the more agents that perform the task, the higher the probability of achieving the task's objective. Furthermore, for the task in question, there are multiple agents in the environment. In the control method, based on observation information related to the agent, other agents in the vicinity of the agent, and the task, request parameters related to whether to request support and response parameters related to whether to respond to requests from other agents are calculated. The process involves calculating the importance of each task associated with a given agent based at least on the request parameters of other agents and the response parameters of that agent. The importance of the task corresponding to the observation information associated with that agent is calculated based on the target value of the task importance corresponding to the observation information, which is output from the policy after inputting the observation information into the policy for each agent. If the response parameter indicates that no response is needed from other agents, the importance of the task being performed by that agent, corresponding to the observation information associated with that agent, is calculated in a manner that brings the importance of the task to be performed closer to the target value. When the response parameter indicates a response to a request from another agent, and the request parameter of the other agent indicates a request for support, the importance of the task corresponding to the observation information associated with that agent is calculated based on the difference between the importance of the task being performed by the other agent and the importance associated with that agent for that task, in a way that increases the importance of the task being performed by the other agent. The task that the agent should perform is selected based on the importance level. Control is performed in a manner that causes the agent to perform the selected task.

8. A computer-readable medium having a program stored thereon that implements a control method for controlling an intelligent agent performing a task, wherein, For the task in question, the more agents that perform the task, the higher the probability of achieving the task's objective. For the task in question, there are multiple agents in the environment... The program causes the computer to perform the following steps: The steps involve calculating request parameters related to whether to request support and response parameters related to whether to respond to requests from other agents, based on observation information related to the agent, other agents in the vicinity of the agent, and the task. The process involves calculating the importance of each task associated with a given agent based at least on the request parameters and response parameters of other agents. This includes: calculating the importance of the task associated with the observation information based on a target value of the task's importance corresponding to the observation information, inputting the observation information into a policy learned by each agent and outputting that policy; calculating the importance of the task associated with the observation information in a manner that brings the importance of the task being performed by the agent closer to the target value when the response parameters indicate no response to a request from another agent; and calculating the importance of the task associated with the observation information in a manner that increases the importance of the task being performed by the other agent based on the difference between the importance of the task being performed by the other agent and the agent-related importance of that task when the response parameters indicate a response to a request from another agent and the request parameters of the other agent indicate a request for support. The steps for selecting the task that the agent should perform are determined based on the importance level. The steps involve controlling the agent to perform the selected task in a manner that enables it to perform the selected task.

Citation Information

Patent Citations

  • Mobile agents for manipulating, moving, and / or reorienting components

    JP2017094122A

  • Robot Control Apparatus

    US20080109114A1

  • Multi-agent reinforcement learning with matchmaking policies

    US20200244707A1

  • Control device, control method, and control system

    WO2019058694A1

  • Methods and apparatus for implementing reinforcement learning

    WO2022152404A1