Information processing device, information processing method, and information processing program
The information processing apparatus addresses inefficiencies in human-robot collaboration by using a hierarchical relational network to automatically adjust device parameters, ensuring efficient operation in response to human input.
Patent Information
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- OMRON CORP
- Filing Date
- 2025-11-04
- Publication Date
- 2026-05-21
AI Technical Summary
Existing systems face inefficiencies in adjusting parameters for controlled devices like robots to operate at a desired level when collaborating with humans, making it difficult to achieve the desired operation without manual and repetitive parameter adjustments.
An information processing apparatus and method that utilizes a relational network with hierarchical elements, including policy and operation elements, to classify and adjust settings based on human input, ensuring consistency and efficiency in device operation.
Enables controlled devices to operate at a desired level by automatically adjusting parameters in response to human input, reducing the need for manual trial and error and enhancing operational efficiency.
Smart Images

Figure JP2025038591_21052026_PF_FP_ABST
Abstract
Description
Information Processing Apparatus, Information Processing Method, and Information Processing Program
[0001] The disclosed technology relates to an information processing apparatus, an information processing method, and an information processing program.
[0002] Non-Patent Document 1 (Mukun Cao et al., 2015, “Automated negotiation for e-commerce decision making: a goal deliberated agent architecture for multi-strategy selection”) discloses an agent architecture that hierarchically has an action plan for a robot and determines the next action from inputs from a user and strategies held.
[0003] Non-Patent Document 2 (Negin Amirshirzad et al., 2019, “Human Adaptation to Human-Robot Shared Control”) discloses a collaborative system between a human and a robot that, based on a control signal from a human, infers a goal point that the human is aiming for as a result of control and generates a robot control signal based on the inference result.
[0004] When a controlled device such as a robot collaborates with another person such as a human, there is a problem that the other person needs to perform detailed and enormous parameter adjustments regarding the behavior of the controlled device, which is inefficient. There is also a problem that it is difficult to perform parameter adjustments at the level desired by the other person and to operate the controlled device at the desired level.
[0005] The disclosed technology has been made in view of the above points, and an object thereof is to provide an information processing apparatus, an information processing method, and an information processing program that can operate a controlled device at a desired level when the controlled device collaborates with another person.
[0006] The information processing device according to the first aspect of the disclosure includes at least a plurality of operation elements for the operation of a controlled device and a plurality of policy or plan elements for a policy or plan, each element having a hierarchical relationship in which it is directly superior or directly inferior to at least one other element, the hierarchical relationship includes information about the strength of the relationship, for policy or plan elements and operation elements that are directly superior or inferior to each other, the policy or plan element is positioned on the superior side of the operation element, the policy or plan element is positioned at the top of the chain of elements connected by a hierarchical relationship, and the operation element is positioned at the bottom of the chain of elements connected by a hierarchical relationship, and the relationship network defines the respective hierarchical relationships between the plurality of elements, and further, the relationship network has a specific hierarchy including a plurality of policy or plan elements that are not in a hierarchical relationship with each other, the specific hierarchy including one set policy or plan element and one set operation element that is directly subordinate to the one set policy or plan element according to the strength of the relationship, or directly subordinate to the one set policy or plan element according to the strength of the relationship The system includes a storage unit that stores a relationship network in which a series of setting elements are pre-configured, each being a chain of setting elements that correspond to a given setting element, with the lowest-level element in the chain being a setting operation element, and which are linked vertically according to the strength of the relationship; a classification unit that classifies information acquired from an external source into corresponding elements in the relationship network; an extraction process that extracts one input policy plan element that belongs to a specific hierarchy among one element that is directly associated to a higher level based on the strength of the relationship, or a chain of one element that is directly associated to a higher level, from the classified elements classified by the classification unit; a determination process that determines whether the currently configured setting policy plan element of a specific hierarchy matches the extracted input policy plan element; and if the determination process determines that they do not match, a confirmation output that confirms whether to change the setting to the input policy plan element.
[0007] The information processing device according to the second aspect of the disclosure includes at least a plurality of operation elements for the operation of a controlled device and a plurality of policy or plan elements for a policy or plan, each element having a hierarchical relationship in which it is directly superior or directly inferior to at least one other element, the hierarchical relationship includes information regarding the strength of the relationship, and for policy or plan elements and operation elements that are directly in a hierarchical relationship, the policy or plan element is positioned on the superior side of the operation element, the policy or plan element is positioned at the top of the chain of elements connected by a hierarchical relationship, and the operation element is positioned at the bottom of the chain of elements connected by a hierarchical relationship, and the relationship network defines the respective hierarchical relationships between the plurality of elements, wherein for one top-level set policy or plan element, there is one set operation element that is directly associated with a lower level according to the strength of the relationship, or a chain of one set element that is directly associated with a lower level according to the strength of the relationship, and the lowest element of the chain is The system includes: a storage unit that stores a relationship network in which a series of setting elements, which are setting operation elements, are linked vertically according to the strength of their relationships and are pre-configured; a classification unit that classifies information acquired from the outside into corresponding elements in the relationship network; an extraction process that extracts the classified elements classified by the classification unit, and one or more elements that are directly or chained to higher levels based on the strength of their relationships with the classified elements, and a series of input corresponding elements that are linked vertically according to the strength of their relationships down to the highest-level policy planning element; a determination process that determines whether there are any elements that are consistent between the series of input corresponding elements and the series of setting elements; and, if the determination process determines that there are no elements that are consistent, a processing unit that outputs a confirmation output to confirm whether to change the setting to one of the input corresponding elements in the series of input corresponding elements.
[0008] The information processing method relating to the third aspect of the disclosure is an information processing method in which a computer classifies information acquired from an external source into corresponding elements in a relational network, extracts one input policy / plan element that belongs to a specific hierarchy among one element that is directly associated with a higher level based on the strength of the relationship, or a chain of one element that is directly associated with a higher level, for the classified elements, performs an extraction process to determine whether the currently set policy / plan element, which is an element of a specific hierarchy, and the extracted input policy / plan element are consistent, and if the determination process determines that they are not consistent, outputs a confirmation output to confirm that the setting should be changed to the input policy / plan element, wherein the relational network includes at least a plurality of operation elements for the operation of a controlled device and a plurality of policy / plan elements for policies or plans, and each element is directly associated with at least one other element or A relational network has a direct hierarchical relationship, where the hierarchical relationship includes information about the strength of the relationship, and for policy planning elements and action elements that are directly in a hierarchical relationship, the policy planning element is positioned higher than the action element, the policy planning element is positioned at the top of the chain of elements connected in a hierarchical relationship, and the action element is positioned at the bottom of the chain of elements connected in a hierarchical relationship, thus defining the hierarchical relationships between multiple elements, and furthermore, a relational network having a specific hierarchy that includes multiple policy planning elements that are not in a hierarchical relationship with each other, wherein one set policy planning element included in the specific hierarchy and one set action element that is directly associated with the one set policy planning element in accordance with the strength of the relationship, or a chain of one set element that is directly associated with the strength of the relationship, where the lowest element in the chain is the set action element, and a series of set elements that are linked vertically in a hierarchical relationship in accordance with the strength of the relationship are pre-configured in this relational network.
[0009] The information processing method relating to the fourth aspect of the disclosure is an information processing method in which a computer performs an extraction process to extract a series of input corresponding elements that are directly or chained to a higher level based on the strength of the relationship with the classified elements, up to the highest-level policy or plan element, based on the strength of the relationship with the classified elements, and a determination process to determine whether there are any elements that are consistent between the series of input corresponding elements and the series of setting elements, and if the determination process determines that there are no elements that are consistent, the computer performs a process to output a confirmation output to confirm that the setting should be changed to one of the input corresponding elements in the series of input corresponding elements, wherein the relationship network comprises a plurality of operation elements for the operation of a controlled device and a plurality of elements for policies or plans A relational network that defines the respective hierarchical relationships between multiple elements, wherein each element has a hierarchical relationship with at least one other element, where each element is directly superior or directly inferior, and the hierarchical relationship includes information about the strength of the relationship, and for hierarchical elements and operational elements that are directly superior to each other, the hierarchical element is positioned on the superior side of the operational element, the hierarchical element is positioned at the top of the chain of elements connected by a hierarchical relationship, and the operational element is positioned at the bottom of the chain of elements connected by a hierarchical relationship, and a relational network that defines the respective hierarchical relationships between multiple elements, wherein for one top-level configured configured hierarchical element, there is one configured operational element that is directly associated with a lower level according to the strength of the relationship, or a chain of one configured element that is directly associated with a lower level according to the strength of the relationship, and the lowest element in the chain is a configured operational element, and a series of configured elements that are linked vertically according to the strength of the relationship are pre-configured.
[0010] The information processing program according to the fifth aspect of the disclosure is an information processing program that causes a computer to perform the following processes: classify information obtained from an external source into corresponding elements in a relational network; extract one input policy / plan element that belongs to a specific hierarchy among one element that is directly associated with a higher level based on the strength of the relationship, or a chain of one element that is directly associated with a higher level, from the classified elements; perform a determination process to determine whether the currently set policy / plan element, which is an element of a specific hierarchy, and the extracted input policy / plan element are consistent; and if the determination process determines that they are not consistent, output a confirmation output to confirm that the setting should be changed to the input policy / plan element, wherein the relational network includes at least a plurality of operation elements for the operation of a controlled device and a plurality of policy / plan elements for policies or plans, and each element is directly associated with at least one other element. A relational network having a hierarchy where one element is superior to the other or directly subordinate to the other, where the hierarchy includes information about the strength of the relationship, where the policy / planning element is positioned higher than the action element in a direct hierarchical relationship, the policy / planning element is positioned at the top of a chain of elements connected by a hierarchical relationship, and the action element is positioned at the bottom of a chain of elements connected by a hierarchical relationship, defining the hierarchical relationships between multiple elements, and further, a relational network having a specific hierarchy that includes multiple policy / planning elements that are not in a hierarchical relationship with each other, wherein one set policy / planning element included in the specific hierarchy and one set action element that is directly subordinate to the one set policy / planning element according to the strength of the relationship, or a chain of one set element that is directly subordinate to the one set policy / planning element according to the strength of the relationship, where the lowest element in the chain is the set action element, and a series of set elements that are linked vertically according to the strength of the relationship are pre-configured in the relational network.
[0011] The information processing program relating to the sixth aspect of the disclosure is an information processing program which causes a computer to perform the following: a classification unit that classifies information acquired from an external source into corresponding elements in a relational network; an extraction process that extracts the classified elements classified by the classification unit and one or more elements that are directly or chainedly associated with the classified elements at a higher level based on the strength of their relationship, and which are chained based on the strength of their relationship down to the highest-level policy or plan element; a determination process that determines whether there are any elements that are consistent between the series of input corresponding elements and the series of setting elements; and if the determination process determines that there are no elements that are consistent, an output that confirms whether the setting should be changed to one of the input corresponding elements in the series of input corresponding elements, wherein the relational network comprises a plurality of operation elements for the operation of a controlled device and a policy or plan. A relational network that defines the respective hierarchical relationships between multiple elements, wherein each element has a hierarchical relationship with at least one other element, where each element is directly superior or directly inferior, and the hierarchical relationship includes information about the strength of the relationship, and for policy and plan elements and action elements that are directly superior to each other, the policy and plan element is positioned on the superior side of the action element, the policy and plan element is positioned at the top of the chain of elements connected by a hierarchical relationship, and the action element is positioned at the bottom of the chain of elements connected by a hierarchical relationship, and a relational network that defines the respective hierarchical relationships between multiple elements, wherein for one top-level set policy and plan element, there is one set action element that is directly associated with a lower level according to the strength of the relationship, or a chain of one set element that is directly associated with a lower level according to the strength of the relationship, and the lowest element in the chain is a set action element, and a series of set elements that are linked vertically according to the strength of the relationship are pre-configured.
[0012] According to the disclosed technology, when a controlled device collaborates with others, it can be made to operate at a desired level.
[0013] This is a configuration diagram showing the hardware configuration of the robot control device according to the first embodiment. This is a configuration diagram showing the functional configuration of the robot control device according to the first embodiment. This is a flowchart of the robot control process according to the first embodiment. This is a diagram showing an example of a confirmation requirement table. This is a configuration diagram showing the functional configuration of the robot control device according to the second embodiment. This is a diagram showing an example of proposal appropriateness. This is a configuration diagram showing the functional configuration of the robot control device according to the third embodiment. This is a diagram showing an example of proposal appropriateness table data. This is a configuration diagram showing the functional configuration of the robot control device according to the fourth embodiment. This is a flowchart of the interaction execution process. This is a diagram showing an example of transmission means table data. This is a diagram showing an example of task achievement condition table data. This is a configuration diagram showing the functional configuration of the robot control device according to the fifth embodiment. This is a flowchart of the robot control process according to the fifth embodiment. This is a diagram showing an example of immediate execution judgment table data. This is a configuration diagram showing the functional configuration of the robot control device according to the sixth embodiment. This is a diagram showing an example of proposal appropriateness table data. This is a diagram showing an example of a relationship network as an action plan hierarchical structure according to the seventh embodiment. This is a diagram showing an example of a relationship network as an action plan hierarchical structure according to the seventh embodiment. This is a diagram showing an example of a relationship network as an action plan hierarchical structure according to the eighth embodiment. This is a diagram showing an example of a relationship network as an action plan hierarchical structure according to the eighth embodiment. This is a flowchart of the robot control process according to the eighth embodiment. This figure shows an example of a relational network as an action plan hierarchical structure according to the 8th embodiment. This figure shows an example of a relational network as an action plan hierarchical structure according to the 9th embodiment. This is a flowchart of the robot control process according to the 9th embodiment. This figure shows an example of a confirmation requirement table according to the 9th embodiment. This figure shows an example of a confirmation requirement table according to the 10th embodiment.
[0014] Hereinafter, an example of an embodiment of the present disclosure will be described with reference to the drawings. In each drawing, identical or equivalent components and parts are given the same reference numerals.
[0015] In the following, "Action Plan" corresponds to "Setting Policy Planning Element" or "Setting Action Element." Furthermore, "Draft Action Plan" in the following corresponds to "Input Policy Planning Element" or "Input Action Element."
[0016] <First Embodiment>
[0017] Figure 1 is a block diagram showing the hardware configuration of the robot control device 10 according to this embodiment. As shown in Figure 1, the robot control device 10 includes a controller 11. The controller 11 is composed of a device including a general-purpose computer.
[0018] As shown in Figure 1, the controller 11 comprises a CPU (Central Processing Unit) 11A, a ROM (Read Only Memory) 11B, a RAM (Random Access Memory) 11C, and an input / output interface (I / O) 11D. The CPU 11A, ROM 11B, RAM 11C, and I / O 11D are connected to each other via a bus 11E. The bus 11E includes a control bus, an address bus, and a data bus.
[0019] Furthermore, the communication unit 12 and the storage unit 13 are connected to the I / O 11D.
[0020] The communication unit 12 is an interface for data communication with the robot RB, behavior sensor BS, etc., as shown in Figure 2.
[0021] The memory unit 13 is composed of a non-volatile external storage device such as a hard disk. As shown in Figure 1, the memory unit 13 stores the robot control program 13A, the learned classification model 13B, the action plan model 13C, and the confirmation requirement table 13D, etc.
[0022] CPU 11A is an example of a computer. Here, "computer" refers to a processor in a broad sense, including general-purpose processors (e.g., CPUs) or specialized processors (e.g., GPUs: Graphics Processing Units, ASICs: Application Specific Integrated Circuits, FPGAs: Field Programmable Gate Arrays, programmable logic devices, etc.).
[0023] The robot control program 13A may be stored in the storage unit 13 by being stored on a non-volatile, non-transitory recording medium, or by being distributed via a network and appropriately installed on the robot control device 10.
[0024] Examples of non-volatile, non-transition recording media include CD-ROMs (Compact Disc Read Only Memory), magneto-optical disks, HDDs (hard disk drives), DVD-ROMs (Digital Versatile Disc Read Only Memory), flash memory, and memory cards.
[0025] The robot control device 10 may be installed on the robot RB, or it may be a separate and independent device from the robot RB. The robot control device 10 is an example of an information processing device of the present disclosure. In this embodiment, the case in which the robot control device 10 controls the robot RB will be described, but the controlled device that is the controlled object of the information processing device of the present disclosure is not limited to a robot.
[0026] In this embodiment, as an example, we will describe a case where the robot control device 10 is a so-called fleet manager, and the robot RB is an autonomous mobile robot (AMR) equipped with an audio output device and a display, etc.
[0027] Figure 2 is a block diagram showing the functional configuration of the robot control device 10. As shown in Figure 2, the robot control device 10 functionally comprises a classification unit 20 and a processing unit 30.
[0028] The CPU 11A functions as one of the functional units shown in Figure 2 by reading and executing the robot control program 13A stored in the memory unit 13.
[0029] The classification unit 20 classifies information acquired from external sources, such as the behavior sensor BS, into corresponding layers and elements within the relational network. The information acquired from external sources input to the classification unit 20 does not have to be language; for example, it may be human behavior, facial expressions, or characters entered via keyboard. Furthermore, the classification unit 20 may choose not to classify information if it does not correspond to any of the elements.
[0030] A relational network is a relational network that includes at least multiple motion elements for the operation of a robot RB as an example of a controlled device, and multiple policy or plan elements for a policy or plan, wherein each element has a hierarchical relationship with at least one other element, which is either directly superior or directly inferior, and the hierarchical relationship includes information about the strength of the relationship, and for policy or plan elements and motion elements that are directly in a hierarchical relationship, the policy or plan element is positioned on the superior side of the motion element, the policy or plan element is positioned at the top of the chain of elements connected by a hierarchical relationship, and the motion element is positioned at the bottom of the chain of elements connected by a hierarchical relationship, thereby defining the respective hierarchical relationships between multiple elements, and furthermore, a relational network having a specific hierarchy that includes multiple policy or plan elements that are not in a hierarchical relationship with each other, wherein the specific hierarchy includes one set policy or plan element, and one set motion element that is directly associated with the set policy or plan element according to the strength of the relationship, or a chain of one set element that is directly associated with the strength of the relationship, where the lowest element in the chain is the set motion element, and a series of set elements that are linked vertically according to the strength of the relationship are pre-configured in the network.
[0031] The strength of the relationships can be individually quantified and stored, or calculated using a formula or similar method. Furthermore, the strength of the relationships can be learned and stored.
[0032] Furthermore, elements included in a particular hierarchy may be top-level elements, or they may be intermediate policy and planning elements rather than top-level elements.
[0033] Furthermore, the relational network includes a first layer containing multiple operational elements for operational content, a second layer containing multiple policy / plan elements for policies or plans that are higher-level than the operational content of the first layer, and a third layer containing multiple higher-level policy / plan elements for higher-level policies or plans that are higher-level than the policies and plans of the second layer. A specific layer may be the second or third layer. Note that higher-level policies or plans may include goals or objectives.
[0034] The processing unit 30 receives the hierarchy and elements provided by the classification unit 20, calculates the difference between the strength of the relationship between the input element and the element currently set in a higher hierarchy than the input hierarchy, and the difference between the strength of the relationship between the element currently set in the input hierarchy and the element in a higher hierarchy than the currently set hierarchy, determines whether the calculated difference exceeds a threshold, and if it is determined that it exceeds the threshold, outputs a confirmation output asking for confirmation on changing the setting to an input policy planning element.
[0035] In this embodiment, we will describe a case where the relational network as an action plan hierarchy structure includes three layers from top to bottom: a policy element (Goal) for the robot RB's policy, a planning element (Plan) for the robot RB's plan, and an action element (Action) for the robot RB. In this embodiment, Action corresponds to the first layer, Plan to the second layer, and Goal to the third layer.
[0036] The classification unit 20 receives messages from the behavior sensor BS, which detects the actions of agent AG, another entity, relative to robot RB. Here, agent AG may be a person or a robot. The behavior sensor BS includes sensors that, for example, detect speech uttered by agent AG or movements such as gestures performed by agent AG. The classification unit 20 may also receive messages from agent AG via keyboard input or data communication.
[0037] The classification unit 20 classifies the messages from the agent AG into any of the action plan hierarchies and elements of Goal, Plan, and Action. Specifically, the classification unit 20 classifies the messages of the agent AG for the robot RB into any of the action plan hierarchies and elements using the learned classification model 13B. Note that the action plan hierarchy in this embodiment includes three hierarchies: a hierarchy including the above-described policy elements, a hierarchy including plan elements, and a hierarchy including action elements.
[0038] The learned classification model 13B is a model that takes the message of the agent AG as input data and outputs, as output data, to which action plan hierarchy and element of Goal, Plan, and Action the input message corresponds.
[0039] For example, suppose the instruction for the robot RB is "reduce the margin" for another robot. In this case, the learned classification model 13B outputs Plan: [margin, -], which represents the plan element of reducing the margin, as output data.
[0040] The processing unit 30 sets the action policy of the robot RB based on the input policy plan element obtained by inputting the classified elements classified by the classification unit 20 into the action plan model 13C that has learned the mutual relationship of the three action plan hierarchies and elements of Goal, Plan, and Action.
[0041] The action plan model 13C is a model that has learned the mutual relationship of each action plan hierarchy and element of Goal, Plan, and Action for the target task. For example, when a specific Plan is input as input data, it outputs, as output data, the likelihood of the action policy plans of Goal and Action.
[0042] Specifically, for example, when the input data represents a Plan: [margin, -] that reduces the margin, as the output data, when the element of Goal is "A", the likelihood is 0.7; when the content of Goal is "B", the likelihood is 0.2; when the element of Goal is "C", the likelihood is 0.1; when the element of Action is "A", the likelihood is 0.3; when the element of Action is "B", the likelihood is 0.2; when the element of Action is "C", the likelihood is 0.1. The following output data is output.
[0043] [Goal: [A]=0.7, Goal: [B]=0.2, Goal: [C]=0.1], [Action: [A]=0.3, Action: [B]=0.2, Action: [C]=0.5])
[0044] In addition, the processing unit 30 calculates the difference between the strength of the relationship between the input element and the element currently set in the upper layer of the input hierarchy, and the strength of the relationship between the element currently set in the input hierarchy and the element in the upper layer of the currently set hierarchy. It determines whether the calculated difference exceeds the threshold value. If it is determined that the calculated difference exceeds the threshold value, a message as a confirmation output for confirming whether to change the setting to the input policy planning element is output to the agent AG.
[0045] Next, the robot control process executed by the CPU 11A of the robot control device 10 will be described with reference to the flowchart shown in FIG. 3. The process in FIG. 3 is repeatedly executed.
[0046] In step S100, the CPU 11A inputs the voice input signal of the agent AG detected by the action sensor BS into the learned classification model 13B, and acquires the action plan layer and elements corresponding to the voice input signal among the plurality of action plan layers included in the relationship network as the action plan hierarchy structure.
[0047] In step S101, the CPU 11A updates the input policy planning elements as a proposed action plan. Specifically, it inputs the classified elements obtained in step S100 into the action plan model 13C and obtains other action plan hierarchy and elements that have the highest likelihood.
[0048] In step S102, the CPU 11A determines whether the input policy planning elements, which are the proposed action plan updated in step S101, deviate from (are consistent with) the elements currently set as action plans. That is, it calculates the difference between the input policy planning elements, which are the proposed action plan updated in step S101, and the elements currently set as action plans.
[0049] In step S103, the CPU 11A determines, based on the calculation result of step S102, whether the input policy planning elements, which are the proposed action plan updated in step S101, deviate from the currently set policy planning elements, i.e., whether there is a difference. If the determination in step S103 is affirmative, the process proceeds to step S104; if the determination in step S103 is negative, the routine terminates.
[0050] In step S104, the CPU 11A determines whether interaction is necessary. That is, it determines whether it is necessary to confirm the input policy plan elements, which are the proposed action plan updated in step S103, with agent AG.
[0051] Specifically, the confirmation requirement table 13D is referenced to determine whether the input policy plan element, which is the proposed action plan updated in step S103, is an input policy plan element that needs to be confirmed with agent AG. The confirmation requirement table 13D will be described later.
[0052] In step S105, based on the result of the interaction necessity determination in step S104, it is determined whether interaction is necessary, that is, whether confirmation from agent AG is necessary. If it is determined that confirmation from agent AG is necessary, the process proceeds to step S106; if it is determined that confirmation from agent AG is not necessary, the process proceeds to step S107.
[0053] In step S106, the CPU 11A performs an interaction. Specifically, it instructs the robot RB to output, for example, the input policy planning elements, which were updated in step S101 as proposed action plans, as voice.
[0054] In step S107, the CPU 11A adopts the input policy planning elements as the proposed action plan updated in step S101 and sets them as the action plan.
[0055] In step S108, the CPU 11A determines the inconsistency between layers. Specifically, it determines whether the changes in the input policy plan elements, which are proposed action plans updated in step S101, are inconsistent with the higher layer compared to the currently set action plan elements. Specifically, the input policy plan elements, which are proposed action plans updated in step S101, are input into the action plan model 13C. As a result, the likelihood of the currently set action plan elements in the higher layer for the input policy plan elements is obtained as output data from the action plan model 13C. If the likelihood of the current action plan elements in the higher layer for the input policy plan elements for the input policy plan elements input into the action plan model 13C is below a threshold, it is determined that an inconsistency has occurred. On the other hand, if the likelihood of the current action plan elements in the higher layer for the input policy plan elements for the input policy plan elements input into the action plan model 13C exceeds a threshold, it is determined that no inconsistency has occurred.
[0056] In step S109, the CPU 11A determines whether or not it determined in step S108 that an inconsistency occurred between layers. If it determines that an inconsistency has occurred, it proceeds to step S110; if it determines that no inconsistency has occurred, this routine terminates.
[0057] In step S110, the CPU 11A inputs the input policy planning elements as action plans updated in step S107 into the action plan model 13C, and sets the element with the highest likelihood of being an action plan proposal element in the higher layer from the output data output from the action plan model 13C as the input policy planning element for the action plan proposal. Then, the process proceeds to step S102 and the same process is repeated.
[0058] The following describes a specific example of the flowchart in Figure 3 where the process flows in the order of steps S100-105 and S107-S110. Furthermore, it is assumed that Agent AG is a human. It is also assumed that Robot RB's Goal is set to "stability," Robot RB's Plan is set to "increase margin" with other robots, and Robot RB is performing Actions corresponding to the set Goal and Plan. Finally, the case where the message from Agent AG is voice will be described.
[0059] In step S100, the CPU 11A inputs the voice input signal of agent AG detected by the behavior sensor BS to the trained classification model 13B and obtains the behavior plan hierarchy and elements corresponding to the voice input signal from among the multiple behavior plan hierarchys included in the relational network as a behavior plan hierarchy structure. Here, let's assume that the voice input signal was an instruction from a person to "reduce the margin". In this case, the trained classification model 13B outputs Plan: [margin, -] which represents the plan element to reduce the margin, as a classified element.
[0060] In step S101, the CPU 11A updates the input policy planning elements as proposed action plans. Here, Plan: [margin, -] is input to the action plan model 13C, and the one with the highest likelihood among the output data output by the action plan model 13C ([Goal: [Efficiency], Action: [margin, 800]) is selected as the input policy planning element as a proposed action plan. This input policy planning element as a proposed action plan means that the Goal element is set to efficiency and the Action element is set to a margin of 800 mm.
[0061] In step S102, the CPU 11A calculates whether the input policy planning element, which is the proposed action plan updated in step S101, deviates from the currently set policy planning element. That is, it calculates the difference between the input policy planning element, which is the proposed action plan updated in step S101, and the currently set policy planning element. Here, we assume that the currently set policy planning element Plan is Plan:[margin,+] and the input policy planning element, which is the proposed action plan updated in step S101, is Plan:[margin,-]. In this case, the difference is Plan:[margin,-].
[0062] In step S103, the CPU 11A determines, based on the calculation results of step S102, whether there is a difference between the input policy planning elements, which are the proposed action plan updated in step S101, and the currently set policy planning elements. Since there is a difference, the process proceeds to step S104.
[0063] In step S104, the CPU 11A determines whether interaction is necessary. That is, it determines whether the input policy planning elements, which are the proposed action plan updated in step S103, need to be confirmed with agent AG. Specifically, it refers to the confirmation requirement table 13D and determines whether the input policy planning elements, which are the proposed action plan updated in step S103, are input policy planning elements that need to be confirmed with a person.
[0064] The Confirmation Requirement Table 13D is a table that specifies whether or not confirmation from Agent AG is required when a difference occurs between the input policy plan elements as a proposed action plan and the elements currently set as the action plan.
[0065] Specifically, as shown in Figure 4, the confirmation requirement table 13D is a table that shows the correspondence between the input policy plan elements as the updated action plan proposal, the changes from the elements as the currently set action plan, the reasons for the changes, the content of the changes, the magnitude of the changes, and whether confirmation is required.
[0066] For example, in the above example, the change is in "Plan", the reason for the change is "reflecting human instructions", the change is in "margin", and the change amount is "-", so confirmation is not required.
[0067] In step S105, based on the result of the interaction requirement determination in step S104, it is determined whether interaction is necessary, that is, whether confirmation from a person is required. In this case, confirmation from a person is not required, so the process proceeds to step S107.
[0068] In step S107, the CPU 11A adopts the input policy plan elements as proposed action plans updated in step S101 and sets them as elements of the action plan. Here, [Goal: [Efficiency], Action: [margin, 800] are set as elements of the action plan.
[0069] In step S108, the CPU 11A determines the inconsistency between layers. Here, it is assumed that an inconsistency has occurred because the likelihood of the elements in the higher layer that are current action plan proposals for the input action plan elements that have been input into the action plan model 13C is below a threshold.
[0070] In step S109, the CPU 11A determines whether or not it determined in step S108 that an inconsistency occurred between layers. Here, it is determined that an inconsistency has occurred, so the process proceeds to step S110.
[0071] In step S110, the CPU 11A inputs the input policy planning elements as action plans updated in step S107 into the action plan model 13C, and sets the element with the highest likelihood of being an action plan proposal at a higher level from the output data output from the action plan model 13C as the input policy planning element for the action plan proposal. Here, [Plan: [margin, -]] is input into the action plan model 13C, and (Goal: [Efficiency], Action: [margin, 800]) output from the action plan model 13C is set as the input policy planning element for the action plan proposal. In this way, the content of Goal, which was initially "stability", is changed to "efficiency" by human instruction.
[0072] Next, we will explain a specific example of the case where the process flows in the order of steps S102-S106 and S100-S106 in the flowchart of Figure 3. In the following, we will assume that agent AG is a human. Furthermore, we will assume that the current action policy element of robot RB's Goal is set to "stability," and that in step S101, the action policy element of Goal is set to "efficiency."
[0073] In step S102, the CPU 11A determines whether the input policy planning elements, which are the proposed action plan for Goal updated in step S101, deviate from the currently set policy planning elements, which are the action plan for Goal. Here, it is determined that the input policy planning elements, which are the proposed action plan for Goal updated in step S101, deviate from the currently set policy planning elements, which are the action plan for Goal.
[0074] In step S103, the CPU 11A determines that there is a discrepancy based on the judgment result in step S102, and proceeds to step S104.
[0075] In step S104, the CPU 11A determines whether it needs to confirm with agent AG the "efficiency improvement" element of the proposed action plan for Goal, which was updated in step S101, by referring to the confirmation necessity table 13D. Here, the change is in "Goal", the reason for the change is "robot's spontaneous change (unconfirmed)", the change content is "-", and the change range is "-", so confirmation is required.
[0076] In step S105, based on the result of the interaction requirement determination in step S104, it is determined whether interaction is necessary, that is, whether confirmation from a person is required. In this case, confirmation from a person is required, so the process proceeds to step S106.
[0077] In step S106, the CPU 11A performs an interaction. Here, it communicates with the person by having the robot RB output the voice message, "Do you want to increase efficiency?"
[0078] In step S100, the CPU 11A inputs the voice input signal from the person detected by the behavior sensor BS to the trained classification model 13B and obtains the behavior plan hierarchy and elements corresponding to the voice input signal from among the multiple behavior plan hierarchys included in the relational network as a behavior plan hierarchy structure. Here, let's assume the person says "Yes". In this case, the trained classification model 13B outputs "Efficiency" as the input policy planning element as the proposed behavior plan for Goal. This is because the person gave an affirmative response of "Yes" to the question "Do you want to increase efficiency?" from the robot control device 10, and therefore "Efficiency" is output as the input policy planning element as the proposed behavior plan for Goal, which is the same as the question.
[0079] In step S101, the CPU 11A updates the input policy planning elements as proposed action plans. Here, Goal: [Efficiency Improvement, +] is input to the action plan model 13C, and the one with the highest likelihood among the output data output by the action plan model 13C ([Plan: [Reduction of stopping positions], Action: [Deletion of stopping position A]) is selected as the input policy planning element as a proposed action plan. This input policy planning element as a proposed action plan means that the content of Plan is set to a reduction in stopping positions, and the content of Action is set to the deletion of stopping position A.
[0080] In step S102, the CPU 11A calculates whether the input policy planning elements, which are the proposed action plan updated in step S101, deviate from the elements currently set as the action plan, i.e., the difference. Here, the difference is [Plan, Reduce stopping position, -], [Goal, Improve efficiency, +].
[0081] In step S103, the CPU 11A determines, based on the calculation results of step S102, whether there is a difference between the input policy planning elements, which are the proposed action plan updated in step S101, and the currently set policy planning elements, which are the action plan. Since there is a difference, the process proceeds to step S104.
[0082] In step S104, the CPU 11A determines whether interaction is necessary. That is, it determines whether it is necessary to have a person confirm the input policy planning elements, which are the updated action plan proposal in step S103. Here, the reason for the change in the "Plan" section is "voluntary change (unconfirmed)", the content of the change is "reduction in stopping position", and the change amount is "-", so confirmation is "needed". Also, the reason for the change in the "Goal" section is "voluntary change (confirmed)", the content of the change is "+", and the change amount is "+", so confirmation is "not needed".
[0083] In step S105, based on the result of the interaction requirement determination in step S104, it is determined whether interaction is necessary, that is, whether confirmation from a person is required. In this case, confirmation from Plan is required, so the process proceeds to step S106.
[0084] In step S106, the CPU 11A performs an interaction. Here, it communicates to the person by having the robot RB output the voice message, "How about reducing the stopping position?"
[0085] Conventionally, it was necessary for a person to manually adjust parameters such as avoidance margins when AMRs (Autonomous Mechanisms) passed each other, and to adjust parameters through trial and error to prevent deadlocks or inefficiencies, in order to operate the robot RB (Robot RB) at the desired level. In contrast, in this embodiment, if a person gives rough instructions at the higher level of the action plan hierarchy, the elements of the action policy at the lower level of the action plan hierarchy are adjusted while confirming the person's ideas and checking for inconsistencies with other action plan hierarchy levels. As a result, the robot RB can be operated at the desired level without the need for a person to manually adjust parameters through trial and error.
[0086] <Second Embodiment>
[0087] Next, a second embodiment will be described. Note that the same reference numerals are used for parts identical to those in the first embodiment, and detailed descriptions are omitted.
[0088] Figure 5 is a functional block diagram of the robot control device 10A according to the second embodiment. As shown in Figure 5, the robot control device 10A has a configuration in which a proposal suitability calculation unit 32 is added to the robot control device 10 of Figure 2. The flowchart of the robot control processing executed by the robot control device 10A is basically the same as the flowchart in Figure 3, but differs in that the likelihood of Action in the processing of step S101 is multiplied by the proposal suitability, which will be described later.
[0089] The following describes the case where the robotic RB is a bicycle-type ergometer used to measure exercise load in rehabilitation settings and gyms.
[0090] The proposal suitability calculation unit 32 calculates the suitability of the proposed input policy planning elements or input action elements as proposed action plans for the robot RB extracted by the processing unit 30.
[0091] In this embodiment, the processing unit 30 sets the ergometer intensity, target running time, and rotation speed as training setting conditions, for example, as Actions for input operating elements of the robot RB. Then, the processing unit 30 sets the Actions of the robot RB based on the proposed suitability score calculated by the proposed suitability score calculation unit 32.
[0092] The proposal suitability calculation unit 32 calculates the appropriateness based on the training setting conditions, the number of trials, and the expected success rate.
[0093] The number of trials is the number of times the ergometer was operated under the set conditions. The expected success rate is the probability that a user of the ergometer will successfully complete training under the set conditions.
[0094] The expected success rate is obtained as follows: For example, the success rate for a small number of training conditions, i.e., combinations of intensity, target running time, and cadence, is taken as input, and the expected success rate for all training conditions is obtained.
[0095] For example, using pre-obtained success rate data for all training conditions from beginners to several advanced users as a template, template matching is applied to the current user's success rate, and the template with the greatest similarity is adopted as the estimate. In other words, the expected success rate is obtained by assuming that the user's ability score is the same as that of the person with the most similar performance among the pre-obtained success rate data.
[0096] Then, the Action of the robot RB set by the processing unit 30 is used as the training setting condition i, and the number of trials for the setting condition is... i , target expected success rate, and expected success rate i Based on this, the appropriateness of the proposal i This is calculated using the following formula. Here, the target expected success rate is the desired expected success rate.
[0097] Proposal suitability i = α × (1 / number of trials) i ) + β × | Target expected success rate - Expected success rate i |) (i ∈ Training setting conditions) ... (1)
[0098] Here, α is a parameter that decreases as the total number of trials increases, and β is a parameter that increases as the total number of trials increases.
[0099] Figure 6 shows an example of the relationship between training settings, number of trials, expected success rate, and proposal suitability.
[0100] The processing unit 30 determines the final likelihood by multiplying the likelihood of robot RB's Action by the proposal suitability score. Then, it adopts the Action with the highest likelihood among robot RB's Actions.
[0101] As a result, in the initial stages with a small number of trials, the system operates in a divergent mode that prioritizes suggesting settings that have not been tried before. From the middle stages onward, as the number of trials increases, it switches to a convergent mode that prioritizes suggesting settings that prioritize the expected success rate. Therefore, in divergent mode, the training settings are suggested that gradually lower the target expected success rate and challenge the user with increasingly difficult conditions.
[0102] Furthermore, the expected success rate decreases to a certain value during the divergent mode. However, it also fluctuates when the Plan transitions to "difficult," "easy," etc. Additionally, the divergent and convergent modes fluctuate depending on requests from people, such as requests classified as Plan, like "I want to try various things."
[0103] Traditionally, users and instructors have had to manually adjust detailed parameters and go through trial and error to find the appropriate training load. This can sometimes lead to excessive safety concerns that prevent them from trying sufficient exercise intensity, or conversely, to taking on challenges that are too difficult.
[0104] In contrast, in this embodiment, the training settings are configured to switch from divergent mode to convergent mode, making it possible to perform training according to a medium- to long-term plan.
[0105] <Third Embodiment>
[0106] Next, a third embodiment will be described. Note that the same reference numerals are used for parts identical to those in the first embodiment, and detailed descriptions are omitted.
[0107] Figure 7 is a functional block diagram of the robot control device 10B according to the third embodiment. As shown in Figure 7, the robot control device 10B has a configuration in which a proposal suitability acquisition unit 34 is added to the robot control device 10 of Figure 2. The flowchart of the robot control processing executed by the robot control device 10B is basically the same as the flowchart in Figure 3, but differs in that the likelihood of Goal, Plan, and Action in the processing of step S101 is multiplied by the proposal suitability score, which will be described later.
[0108] In the following description, similar to the first embodiment, we will explain the case where the robot control device 10B is a fleet manager and the robot RB is an autonomous mobile transport robot.
[0109] The proposal suitability acquisition unit 34 acquires a proposal suitability score, weighted by the recommendation scores from multiple perspectives, for input policy plan elements or input action elements as proposed action plans for any of the action plan hierarchies of Goal, Plan, or Action extracted by the processing unit 30.
[0110] Recommendation levels from multiple perspectives include, for example, safety recommendations, administrator recommendations, and predictive database recommendations.
[0111] The safety recommendation score is assigned a higher value to input policy planning elements or input action elements that represent safer proposed action plans. For example, a higher weight is assigned to input policy planning elements or input action elements that represent proposed action plans that can avoid predefined hazardous conditions, such as congestion in passageways.
[0112] The administrator recommendation score is set directly based on the administrator's input. For example, if the administrator inputs a change to an administrator recommendation score of 0.8 when Goal prioritizes efficiency, then 0.8 will be set as the weight.
[0113] The predictive database recommendation score is a general recommendation score for Goal, Plan, and Action (GPA) for a given event, based on pre-prepared data. For example, if the GPA has been updated many times, it is considered that the trial-and-error process is progressing, and a smaller weight is assigned to it. The pre-prepared data is created, for example, based on operational data from various factories.
[0114] Figure 8 shows the proposal suitability table data 40. The proposal suitability table data 40 is table data that shows the correspondence between input policy planning elements or input action elements as proposed action plans for each action plan hierarchy of Goal, Plan, and Action (GPA), events, safety recommendation level, manager recommendation level, prediction DB recommendation level, and proposal suitability level.
[0115] Examples of events include when the environmental sensor ES detects a change in safety-related environmental conditions around the robot RB.
[0116] The appropriateness of the proposal is calculated in advance using a predetermined formula that weights the safety recommendation level, the manager recommendation level, and the prediction database recommendation level.
[0117] The administrator recommendation score is reflected in the proposal suitability table data 40 based on the value directly entered by the administrator, and the proposal suitability score is updated accordingly.
[0118] The proposal suitability acquisition unit 34 acquires the input policy plan element or input action element as a proposed action plan for one of the set Goal, Plan, or Action extracted by the processing unit 30, and the proposal suitability corresponding to the event from the proposal suitability table data 40.
[0119] The processing unit 30 determines the final likelihood by multiplying the likelihood of an input policy planning element or input action element as a proposed action plan (Goal, Plan, or Action) by the proposal suitability score. Then, it adopts the input policy planning element or input action element with the highest likelihood among the various proposed action plans.
[0120] When actually performing a task, the action plan is not determined solely by the user's requests or task efficiency; it is necessary to select the most desirable action plan from multiple perspectives. In this embodiment, multiple perspectives are reflected in the action plan, as well as changes in priorities due to third parties or the surrounding environment. Therefore, the user can perform the task according to an action plan that reflects these changes.
[0121] <Fourth Embodiment>
[0122] Next, a fourth embodiment will be described. Note that the same reference numerals are used for parts identical to those in the first embodiment, and detailed descriptions are omitted.
[0123] Figure 9 is a functional block diagram of the robot control device 10C according to the fourth embodiment. As shown in Figure 7, the robot control device 10C has an interaction adjustment unit 36 added to the robot control device 10 of Figure 2.
[0124] The interaction adjustment unit 36 adjusts at least one of the following for the input policy plan element or input action element, which is an action plan proposal of the action plan hierarchy extracted by the processing unit 30: the timing of outputting the message to be output to agent AG and the adjustment of the message itself.
[0125] Two specific examples will be described below. First, we will describe the case where, similar to the first embodiment, the robot control device 10C is a fleet manager and the robot RB is an autonomous mobile transport robot.
[0126] In this embodiment, in step S106 of Figure 3, the interaction execution process shown in Figure 9 is performed.
[0127] In step S200, the CPU 11A determines the means of message transmission. Specifically, it refers to the transmission means table data 42 shown in Figure 11 and obtains the information transmission means corresponding to the input policy plan element or input action element as an action plan proposal among Goal, Plan, and Action extracted by the processing unit 30. In the example in Figure 11, if the action plan proposal for Goal set by the processing unit 30 is "efficiency-focused," voice is set as the information transmission means, and if the action plan proposal for Action set by the processing unit 30 is "reduce stopping position 4," a screen is set as the information transmission means. Thus, since Goal information can be easily transmitted in language, voice is set as the information transmission means, and since Action information cannot be easily transmitted in language, a screen is set as the information transmission means.
[0128] In step S201, it is determined whether or not an interaction can be performed, for example, whether or not agent AG is speaking. Specifically, if an audio signal is input from the behavior sensor BS, it is determined that agent AG is speaking and the interaction cannot be performed. On the other hand, if no audio signal is input from the behavior sensor BS, it is determined that agent AG is not speaking and the interaction can be performed.
[0129] If it is determined that the interaction can be performed, the process proceeds to step S202. On the other hand, if it is determined that the interaction cannot be performed, the process waits until it is determined that the interaction can be performed.
[0130] In step S202, the CPU 11A performs turn-taking processing. Specifically, for example, it controls the robot RB by turning on an LED on the robot RB to inform agent AG that the robot RB is about to speak. As a result, agent AG recognizes that the robot RB is about to speak.
[0131] In step S203, the CPU 11A performs an interaction, that is, it controls the robot RB to speak a message.
[0132] In this way, the timing of the message output to agent AG is adjusted, preventing robot RB and agent AG from speaking at the same time.
[0133] Next, as another specific example, we will describe the case where the robot control device 10C is applied to a teaching system that facilitates the achievement of an objective. Here, we will explain using the example where the task performed by the user is 3D modeling.
[0134] The following describes the case where the process flows in the order of steps S100 to S106 in the flowchart of Figure 3.
[0135] In step S100, the CPU 11A inputs the voice input signal from the person detected by the behavior sensor BS to the trained classification model 13B and obtains the behavior plan hierarchy and elements corresponding to the voice input signal from among the multiple behavior plan hierarchys included in the relational network as a behavior plan hierarchy structure. Here, let's assume the person says, "I want to learn how to operate it." In this case, the trained classification model 13B outputs the hierarchy and elements as a proposed behavior plan corresponding to "I want to learn how to operate it."
[0136] In step S101, the CPU 11A updates the input policy plan elements or input action elements as proposed action plans. Here, the task achievement condition table data 44 shown in Figure 12 is read, and Actions are selected from the top of the task achievement condition table data 44 in order.
[0137] The task completion condition table data 44, as shown in Figure 12, is table data that represents the correspondence between the user agent AG's task, the user's task completion conditions, and the interaction content. The task includes Goal, Plan, and Action. The interaction content includes the items "Sharing Goal," "Comparison of Completion Conditions Not Met 1," and "Selection Conditions for Not Met 1."
[0138] In this embodiment, the task achievement condition table data 44 is defined such that the user executes Actions sequentially from the top row of the task achievement condition table data 44, and if the user's task achievement condition corresponding to the executed Action is met, the Action in the next row is executed.
[0139] Therefore, in step S102, the CPU 11A calculates the discrepancy between the current task status and the achievement conditions in the task achievement condition table data 44.
[0140] Here, we assume that the current task status is as follows: Goal is set to "Learn basic operations - Improve completion quality," Plan is set to "Create a cube," and Action is set to "System: Display information about the cube to be created." In this case, the user's task completion condition is "Complete the specified cube with an error of 1 [mm] or less on each side."
[0141] Furthermore, let's assume that there is a 2 mm error in the horizontal size of the user's operation. In this case, the CPU 11A determines that there is a discrepancy between the current task status and the achievement conditions in the task achievement condition table data 44.
[0142] In step S103, the CPU 11A determines that there is a discrepancy based on the judgment result in step S102, and proceeds to step S104.
[0143] In step S104, the CPU 11A determines that it needs to confirm with agent AG and proceeds to step S105. In step S105, it determines that interaction is necessary and proceeds to step S106.
[0144] In step S106, the CPU 11A performs an interaction. Here, it refers to the task achievement condition table data 44 and obtains "We are currently aiming to improve the finished quality." from the "Goal Sharing" column. Since the current error corresponds to "The error of the created cube is 2 [mm] or more" in the "Unfulfilled 1 Selection Condition" column, it obtains the interaction content "Almost there!" from the "Achievement Condition Comparison Unfulfilled 1" column and has the robot RB speak it.
[0145] Conventionally, in teaching systems, users perform tasks one by one according to a pre-created plan. As a result, users simply follow the system's instructions, and it can be difficult for them to understand the purpose of their work. In contrast, in this embodiment, the robot RB operates based on multiple objectives and associated action plans, and can appropriately provide feedback to the user regarding the discrepancy between the objective and the current situation. This allows the user to accurately understand the purpose of their current work and what needs to be done next. Furthermore, feedback such as encouragement motivates the user to achieve their objectives, enabling them to achieve their goals efficiently.
[0146] <Fifth Embodiment>
[0147] Next, a fifth embodiment will be described. Note that the same reference numerals are used for parts identical to those in the first embodiment, and detailed descriptions are omitted.
[0148] Figure 13 is a functional block diagram of the robot control device 10D according to the fifth embodiment. As shown in Figure 13, the robot control device 10D has an immediate execution determination unit 38 added to the robot control device 10 of Figure 2.
[0149] The immediate execution determination unit 38, when a difference occurs between the input policy plan element or input action element as a proposed action plan and the currently set policy plan element or input action element as an action plan, immediately determines whether or not to adopt the input policy plan element or input action element as a proposed action plan without sending a confirmation output to the external agent AG regarding the difference.
[0150] Figure 14 is a flowchart of the robot control process executed by the robot control device 10. The robot control process shown in Figure 14 has steps S103A and S103B added to the robot control process in Figure 3. The other processes are the same as those in Figure 3, so their explanation is omitted.
[0151] In step S103A, the CPU 11A determines whether to immediately execute the input policy plan element or input action element as the proposed action plan updated in step S101. Specifically, it refers to the immediate execution determination table data 46 shown in Figure 15 to obtain the necessity of immediate execution according to the event that occurred. If the necessity of immediate execution is "True", it proceeds to step S103B; otherwise, it proceeds to step S104. Situations where immediate execution is necessary include, for example, when the conversation breaks down, such as when agent AG gives an unexpected type of response to robot RB's speech.
[0152] In step S103B, the CPU 11A, similar to step S107, adopts the input policy plan element or input action element as the proposed action plan updated in step S101 and sets it as the set policy plan element or set action element as the action plan. As a result, even if it is determined in step S103 that the input policy plan element or input action element as the proposed action plan deviates from the currently set set policy plan element or set action element as the action plan, the input policy plan element or action element as the proposed action plan can be executed immediately without determining whether confirmation from agent AG is necessary.
[0153] <Sixth Embodiment>
[0154] Next, a sixth embodiment will be described. Note that the same reference numerals are used for parts identical to those in the first embodiment, and detailed descriptions are omitted.
[0155] Figure 16 is a functional block diagram of the robot control device 10E according to the sixth embodiment. As shown in Figure 16, the robot control device 10E has an agent model 39 added to the robot control device 10 of Figure 2.
[0156] In the following section, similar to the second embodiment, we will describe the case where the robot RB is a bicycle-type ergometer used to measure exercise load in rehabilitation settings and gyms.
[0157] The agent model 39 estimates the tendency of the agent's actions as an example of the agent's behavior, based on the action history, which is an example of the agent AG's operation history. The tendency of the agent's actions includes, for example, three elements: recommendation level for change, recommendation level for difficulty, and recommendation level for exploration, each represented by a value of 0 or greater and 1 or less. In this case, the classification unit 20 classifies the actions of the agent AG detected by the action sensor BS into layers such as change, difficulty, and exploration, and inputs them into the agent model 39.
[0158] The processing unit 30 sets input policy planning elements or input action elements as proposed action plans for the robot RB's actions, based on the behavioral tendencies of agent AG estimated by agent model 39. Specifically, as shown in Figure 17, it refers to the pre-prepared proposal suitability table data 48 and obtains proposal suitability corresponding to the input policy planning elements or input action elements as proposed action plans for Goal, Plan, and Action (GPA), as well as the recommendation level for change, recommendation level for difficulty, and recommendation level for exploration estimated by agent model 39. Then, it determines the final likelihood by multiplying the likelihood of the input policy planning elements or input action elements as proposed action plans by the proposal suitability.
[0159] This allows users to not only directly interact with the robot RB's action plan but also to reflect their conceptual intentions. As a result, the robot RB can autonomously formulate and propose action plans based on the user's intentions, enabling it to perform tasks even when the user is unaware of specific objectives or actions.
[0160] <Seventh Embodiment>
[0161] Next, a seventh embodiment will be described. Note that the same reference numerals are used for parts identical to those in the first embodiment, and detailed descriptions are omitted.
[0162] In the seventh embodiment, another example of the action plan model 13C will be described. In the first to sixth embodiments, the action plan model 13C was a model that learned the interrelationships of the Goal, Plan, and Action action plan layers for the target task. In contrast, the relational network as the action plan layer structure in the action plan model 13C according to the seventh embodiment is a tree structure model in which the strength of the connections between upper-level nodes and lower-level nodes is defined, as shown in Figure 18.
[0163] The action planning model 13C shown in Figure 18 is a policy network that determines the actions of the robot RB for the target task, and it has a tree structure. Each node has a defined strength of connection with the connected nodes. When a specific node is specified, this tree-structured action planning model 13C outputs the connected nodes and the strength of the connections.
[0164] In the example in Figure 18, for example, a node defined as "Efficiency" is connected to a subnode defined as "Stopping Position - 1" with a strength of 0.2, and to a subnode defined as "Avoidance Width - 1" with a strength of 0.8. Similarly, a node defined as "Robustness" is connected to a subnode defined as "Stopping Position + 1" with a strength of 0.9, and to a subnode defined as "Avoidance Width - 1" with a strength of 0.1.
[0165] In the example shown in Figure 18, the tree structure has a depth (distance) of two layers, but the depth from the top layer to the bottom layer may vary from node to node. For example, as shown in Figure 19, nodes with depths of two layers and nodes with depths of three layers may be mixed together. Note that the strength of the connections is omitted in Figure 19.
[0166] <Eighth Embodiment>
[0167] Next, the eighth embodiment will be described. Note that the same reference numerals are used for parts identical to those in the first embodiment, and detailed descriptions are omitted.
[0168] In the eighth embodiment, another example of the action plan model 13C is described. In the first to sixth embodiments, the action plan model 13C was a model that learned the interrelationships of the Goal, Plan, and Action layers for the target task. In contrast, in the eighth embodiment, the relational network in the action plan model 13C is a tree structure model in which the strength of the connections between upper-level nodes and lower-level nodes is defined, as shown in Figure 20.
[0169] Here, a relational network is defined as a relational network that includes at least multiple action elements about actions and multiple policy or plan elements about policies or plans, wherein each element has a hierarchical relationship with at least one other element, being either directly superior or directly inferior, and the hierarchical relationship includes information about the strength of the relationship, and for policy or plan elements and action elements that are directly in a hierarchical relationship, the policy or plan element is positioned higher than the action element, the policy or plan element is positioned at the top of the chain of elements connected by a hierarchical relationship, and the action element is positioned at the bottom of the chain of elements connected by a hierarchical relationship, thereby defining the respective hierarchical relationships between multiple elements, and furthermore, a relational network having a specific hierarchy that includes multiple policy or plan elements that are not in a hierarchical relationship with each other, wherein the specific hierarchy includes one set policy or plan element and one set action element that is directly associated with the set policy or plan element according to the strength of the relationship, or a chain of one set element that is directly associated with the strength of the relationship, and the lowest element in the chain is the set action element, and a series of set elements that are linked vertically according to the strength of the relationship are pre-configured in this relational network.
[0170] The strength of the relationships can be individually quantified and stored, or calculated using a formula or similar method. Furthermore, the strength of the relationships can be learned and stored.
[0171] Furthermore, the relational network includes a first layer containing multiple operational elements for operational content, a second layer containing multiple policy / plan elements for policies or plans that are higher-level than the operational content of the first layer, and a third layer containing multiple higher-level policy / plan elements for higher-level policies or plans that are higher-level than the policies and plans of the second layer. A specific layer may be the second or third layer. Note that higher-level policies or plans may include goals or objectives.
[0172] In the example shown in Figure 20, the relational network has nodes T1 and T2 at the top layer, nodes N1 and N2 at the intermediate layer below, nodes N3 to N6 at the intermediate layer below that, and nodes A1 to A4 at the lowest layer below that.
[0173] Node T1 has "efficiency" set as a policy planning element for the policy. Node T2 has "robustness" set as a policy planning element for the policy.
[0174] Node N1 has "Reduce travel distance" set as a policy element for the plan. Node N2 has "Increase travel time" set as a policy element for the plan.
[0175] Node N3 has "Reduce Avoidance Width" set as a policy element for the plan. Node N4 has "Reduce Stopping Position" set as a policy element for the plan. Node N5 has "Increase Number of Allocated Vehicles" set as a policy element for the plan. Node N6 has "Increase Avoidance Width" set as a policy element for the plan.
[0176] Node A1 has "avoidance width 800" set as an action element for its operation. Node A2 has "stopping position - 1" set as an action element for its operation. Node A3 has "number of vehicles + 1" set as an action element for its operation. Node A4 has "avoidance width 1000" set as an action element for its operation.
[0177] In the example in Figure 20, the top layer is designated as a specific layer, but this specific layer could also be an intermediate layer.
[0178] The classification unit 20 classifies information acquired from external sources, such as the behavior sensor BS, into corresponding elements within the relational network. The information acquired from external sources input to the classification unit 20 does not have to be language; for example, it may be human behavior, facial expressions, or characters entered via keyboard. Furthermore, the classification unit 20 may choose not to classify information if it does not correspond to any of the elements.
[0179] The processing unit 30 performs an extraction process to extract one input policy plan element that belongs to a specific hierarchy among one element that can be directly associated with a higher level based on the strength of the relationship, or a chain of one element that can be directly associated with a higher level, from the classified elements classified by the classification unit 20. The processing unit 30 then performs a judgment process to determine whether the extracted input policy plan element is consistent with the currently set setting policy plan element, which is an element of the specific hierarchy. If the judgment process determines that they are not consistent, it outputs a confirmation output asking for confirmation on changing the setting to the input policy plan element.
[0180] Furthermore, the processing unit 30, in response to the confirmation result input for the input confirmation output, resets the input policy plan element as one setting policy plan element included in a specific hierarchy, and resets a series of setting elements consisting of one operation element that is directly associated with the input policy plan element according to the strength of its relationship, or a chain of elements that is directly associated with the input policy plan element according to the strength of its relationship, where the lowest element in the chain is a setting operation element.
[0181] Furthermore, if the processing unit 30 determines in the judgment process that the elements are consistent, and determines that there is a discrepancy between one pre-configured element that is directly associated with a specific hierarchy of one setting policy planning element according to the strength of the relationship, and one element that is directly below the input policy planning element that belongs to the specific hierarchy, among the elements that are directly or chain-linkedly associated with a classified element based on the strength of the relationship, the processing unit 30 may change the settings to an element that is directly or chain-linkedly associated with a classified element based on the strength of the relationship, either higher or lower.
[0182] The robot control process according to this embodiment will be described below with reference to the flowchart shown in Figure 22. Similar to the first embodiment, the description will focus on the case where the robot control device 10C is a fleet manager and the robot RB is an autonomous mobile transport robot. Furthermore, the description will focus on the case where agent AG is a person and messages from agent AG are voiced. Additionally, the "robustness" of node T2 at the top layer is set as a setting policy planning element for robot RB, and each element connected in the order of node T2, N6, and A4 is set as an action policy, as shown by the thick dashed line in Figure 21, based on the initially set likelihood.
[0183] In this situation, in step S200 of Figure 22, the CPU 11A inputs the voice input signal of agent AG detected by the behavior sensor BS to the trained classification model 13B and classifies the elements corresponding to the voice input signal. Here, let's assume the voice input signal was an instruction from a person to "reduce the margin." In this case, the trained classification model 13B outputs "reduce avoidance width" as a classified element, which is node N3 representing the planning element for reducing the margin.
[0184] In step S201, the CPU 11A updates the proposed action plan. Specifically, in order to update both higher and lower levels, it traverses the nodes with the highest probability of being connected to the classified element nodes, and sets each element of the nodes connected from the highest level to the lowest level as the proposed action plan.
[0185] Here, as shown by the thick solid lines in Figure 21, each element connected in the order of nodes T1, N1, N3, and A1, including the classified element node N3, is set as a proposed course of action.
[0186] In step S202, the CPU 11A compares the proposed action plan updated in step S201 with the currently set action plan and determines whether there is a discrepancy (consistency) in the nodes at a specific level. Specifically, it extracts the differences between the proposed action plan updated in step S201 and the currently set action plan, and if there is a difference, it sets the node at the specific level in the proposed action plan updated in step S201 as the discrepancy to be checked. Here, as shown in Figure 20, the node at the specific level in the proposed action plan instructed by agent AG is T1, and the currently set node at the specific level is T2, so there is a difference, and the discrepancy to be checked is set to node T1. If there is no difference, the discrepancy to be checked is not set.
[0187] In step S203, the CPU 11A determines, based on the result of step S202, whether or not there is a discrepancy in a node at a specific level, that is, whether or not a discrepancy target for verification was set in step S202. If a discrepancy target for verification is set, the process proceeds to step S204; otherwise, the routine terminates. In this case, since the discrepancy target for verification is set at node T1, the process proceeds to step S204.
[0188] In step S204, the CPU 11A generates a confirmation message for agent AG as part of the interaction plan, to ask about the elements of the discrepancy to be checked. Here, as a confirmation message to check about the elements of node T1, it generates a message with wording such as "Is the goal to increase efficiency?".
[0189] In step S205, the CPU 11A performs an interaction. Here, it communicates to agent AG by having robot RB output the voice message, "Is the goal to increase efficiency?"
[0190] In step S206, the CPU 11A determines whether the discrepancy at the target discrepancy location has been resolved. That is, it determines whether the response from agent AG detected by the behavior sensor BS was an affirmative response. If the response from agent AG is an affirmative response, the node of the target discrepancy location is registered as a node where the discrepancy has been resolved, and the process returns to step S202. In this case, since node T1 is registered as a node where the discrepancy has been resolved, in step S202, node T1 is excluded from the target discrepancy location. On the other hand, if the response from agent AG is a negative response, the process proceeds to step S207.
[0191] In step S207, the CPU 11A replans the proposed course of action. Specifically, among the nodes of the proposed course of action updated in step S201, the CPU 11A reconfigures each element (a series of setting elements) of the connected nodes that have the highest probability of connecting from the highest level to the lowest level, excluding the nodes that were rejected in step S206, as the next proposed course of action.
[0192] In the above explanation, we described the case where the specific layer is the first layer of the highest layer, but the specific layer is not limited to the highest layer, nor is it limited to the first layer. For example, as shown in Figure 23, the specific layer may be the second intermediate layer from the highest layer. In this case, the process is the same as above, except that in step S202 of Figure 22, the location of the deviation to be checked is set to node N1.
[0193] <Ninth Embodiment>
[0194] Next, the ninth embodiment will be described. Note that the same reference numerals are used for parts identical to those in the first and eighth embodiments, and detailed descriptions are omitted.
[0195] In the ninth embodiment, as shown in Figure 24, we will describe a case where there is no specific layer in the relational network.
[0196] In this embodiment, the processing unit 30 performs an extraction process to extract the classified elements classified by the classification unit 20, and one or more elements that are directly or chain-linked to higher levels based on the strength of their relationship to the classified elements, and a series of input corresponding elements that are chained based on the strength of their relationship down to the highest-level policy / planning element. The processing unit 30 performs a determination process to determine whether there are any elements that match between the series of input corresponding elements and the series of setting elements. If the determination process determines that there are no matching elements, it outputs a confirmation output to confirm whether to change the setting to one of the input corresponding elements in the series of input corresponding elements.
[0197] Furthermore, if the processing unit 30 determines in the judgment process that there are matching elements, it determines whether the lower input corresponding elements directly below the matching element among the series of input corresponding elements and the lower setting elements directly below the matching element among the series of setting elements are matching. If the lower input corresponding elements and the lower setting elements are not matching, it outputs a confirmation output asking for confirmation on changing the settings to part or all of the series of input corresponding elements.
[0198] Furthermore, the processing unit 30 may output a confirmation output if the mismatched element is a policy planning element, but may not output a confirmation output if the mismatched element is an operation element, in cases where the lower input corresponding element and the lower setting element do not match.
[0199] Furthermore, the processing unit 30 changes the settings of a series of input-corresponding elements in accordance with the confirmation result input for the input confirmation output.
[0200] The robot control process related to this embodiment will be explained below with reference to the flowchart shown in Figure 25.
[0201] Furthermore, similar to the eighth embodiment, we will describe the case where the robot control device 10C is a fleet manager and the robot RB is an autonomous mobile transport robot. In addition, we will describe the case where agent AG is a person and the message from agent AG is voice. Furthermore, the "robustness" of node T2 at the top layer is set as a setting policy planning element for robot RB, and each element (a series of setting elements) connected in the order of node T2, N6, and A4 is set as an action policy according to the initial likelihood, as shown by the thick dashed line in Figure 24.
[0202] In this situation, in step S200 of Figure 25, the CPU 11A inputs the voice input signal of agent AG detected by the behavior sensor BS to the trained classification model 13B and classifies the elements corresponding to the voice input signal. Here, let's assume the voice input signal was an instruction from a person to "reduce the margin." In this case, the trained classification model 13B outputs "reduce avoidance width" as a classified element, which is node N3 representing the planning element for reducing the margin.
[0203] In step S201, the CPU 11A updates the proposed action plan. Specifically, in order to update both higher and lower levels, it traverses the nodes with the highest probability of being connected to the classified element nodes, and sets each element (a series of input-corresponding elements) of the nodes connected from the highest level to the lowest level as the proposed action plan.
[0204] Here, as shown by the thick solid lines in Figure 24, each element connected in the order of nodes T1, N1, N3, and A1, including the classified element node N3, is set as a proposed action plan.
[0205] In step S202, the CPU 11A compares the proposed action plan updated in step S201 with the currently set action plan and determines whether there is a discrepancy. Specifically, if all nodes in the proposed action plan updated in step S201 and the currently set action plan match, there is no discrepancy, and no discrepancy location to be checked is set. On the other hand, if all nodes in the proposed action plan updated in step S201 and the currently set action plan do not match, it is determined whether the two contain the same node. If the two contain the same node (a consistent element), the node one level below the identical node in both is set as the discrepancy location to be checked. On the other hand, if the two do not contain the same node, the node at the highest level is set as the discrepancy location to be checked. In this case, the proposed action plan updated in step S201 consists of nodes T1, N1, N3, and A1, and the set action plan consists of nodes T2, N6, and A4, so no identical nodes are included. Therefore, among the series of input correspondence elements in the proposed action plan, for example, node T1 at the highest level is set as the location of the discrepancy to be checked.
[0206] In step S203, the CPU 11A determines, based on the result of step S202, whether or not there is a discrepancy in a node at a specific level, that is, whether or not a discrepancy target for verification was set in step S202. If a discrepancy target for verification is set, the process proceeds to step S204; otherwise, the routine terminates. In this case, since the discrepancy target for verification is set at node T1, the process proceeds to step S204.
[0207] In step S204, the CPU 11A determines whether interaction is necessary. That is, it determines whether it is necessary to confirm with agent AG the elements of the node at the location of the deviation to be checked, as set in step S203.
[0208] Specifically, the need for confirmation is determined by referring to the confirmation requirement table 13D shown in Figure 26. As shown in Figure 26, the confirmation requirement table 13D is a table that shows the correspondence between the action plan draft updated in step S201 and the currently set action plan, the points of change (points of deviation from the target of confirmation), the reason for the change, whether confirmation is required, and the content of confirmation. In the example in Figure 26, the policy planning elements such as nodes T1 and N3 require confirmation, while the operational elements such as nodes A1 and A2 do not require confirmation.
[0209] In this case, the change is at node T1, and the reason for the change is "voluntary change (unconfirmed)," so confirmation is required, and the confirmation content is "Is the goal to increase efficiency?"
[0210] In step S204A, based on the result of the interaction requirement determination in step S204, it is determined whether interaction is necessary, that is, whether confirmation from agent AG is required. In this case, confirmation from agent AG is required, so the process proceeds to step S205.
[0211] In step S205, the CPU 11A performs an interaction. Here, it communicates to agent AG by having robot RB output the voice message, "Is the goal to increase efficiency?"
[0212] In step S206, the CPU 11A determines whether the discrepancy at the target discrepancy location has been resolved. That is, it determines whether the response from agent AG detected by the behavior sensor BS was an affirmative response. If the response from agent AG is an affirmative response, the node of the target discrepancy location is registered as a node where the discrepancy has been resolved, and the process returns to step S202. In this case, since node T1 is registered as a node where the discrepancy has been resolved, in step S202, node T1 is excluded from the target discrepancy location. On the other hand, if the response from agent AG is a negative response, the process proceeds to step S207.
[0213] In step S207, the CPU 11A replans the proposed course of action. Specifically, among the nodes of the course of action updated in step S201, the CPU 11A resets each element of the connected nodes that have the highest probability of connecting from the highest level to the lowest level, excluding the nodes that were rejected in step S206, as the next proposed course of action.
[0214] The following describes an example of changes in the divergence locations to be checked. Here, we assume that in the determination in step S206, all responses are affirmative, and the processes in steps S202 to S206 are executed repeatedly. In this case, the divergence locations to be checked and the registered divergence-resolved nodes in each iteration will be as follows.
[0215] Location of deviation detected in the first check: T1 Nodes where deviation has been resolved: None
[0216] Second check for discrepancies: N1 Node with resolved discrepancy: T1
[0217] Third check for discrepancies: N3 Nodes where discrepancies have been resolved: T1, N1
[0218] Fourth check for discrepancies: A1 Nodes where discrepancies have been resolved: T1, N1, N3
[0219] Fifth check for discrepancies: None. Nodes where discrepancies have been resolved: T1, N1, N3, A1
[0220] Then, if no discrepancies are found in the fifth processing step, the settings are changed to the series of input corresponding elements of nodes T1, N1, N3, and A1.
[0221] <Tenth Embodiment>
[0222] Next, a tenth embodiment will be described. Note that the same reference numerals are used for parts identical to those in the first embodiment, and detailed descriptions are omitted.
[0223] In the tenth embodiment, as an example, the robot control device 10 is a route generation device that generates a route to the destination of the automobile, and the robot RB is a car navigation device equipped with an audio output device and a display, etc., that performs navigation of the route generated by the route generation device.
[0224] In the tenth embodiment, the robot control processing performed by the CPU 11A of the robot control device 10 is the same as the processing in the flowchart shown in Figure 3 described in the first embodiment.
[0225] The following describes a specific example of the process flowing in the order of steps S100-105 and S107-S110 in the flowchart of Figure 3. Furthermore, it is assumed that Agent AG is a human. It is also assumed that Robot RB's Goal is set to "time," representing time saving, and Robot RB's Plan is set to "mountain road +," representing the selection of a mountain path as the route for the purpose of saving time, and that Robot RB is performing Actions corresponding to the set Goal and Plan. Finally, the case where the message from Agent AG is voice will be described.
[0226] In step S100, the CPU 11A inputs the voice input signal of agent AG detected by the behavior sensor BS to the trained classification model 13B and obtains the behavior plan hierarchy and elements corresponding to the voice input signal from among the multiple behavior plan hierarchys included in the relational network as a behavior plan hierarchy structure. Here, let's assume that the voice input signal was an instruction from a person saying "Do not choose the mountain path." In this case, the trained classification model 13B outputs Plan: [mountain, -] which represents the planning element of not choosing the mountain path, as a classified element.
[0227] In step S101, the CPU 11A updates the input policy planning elements as proposed action plans. Here, Plan: [mountain, -] is input to the action planning model 13C, and the input policy planning element with the highest likelihood ([Goal: [time], Action: [route x]) from the output data output by the action planning model 13C is selected as the input policy planning element as the proposed action plan. Note that "time" is an element that means saving time. This input policy planning element as a proposed action plan means that the Goal element is set to time, and the Action element is set to route x (for example, a national highway) away from the mountain road.
[0228] In step S102, the CPU 11A calculates whether the input policy planning element, which is the proposed action plan updated in step S101, deviates from the currently set policy planning element. That is, it calculates the difference between the input policy planning element, which is the proposed action plan updated in step S101, and the currently set policy planning element. Here, let's assume that the currently set policy planning element Plan is Plan: [mountain road, +] and the input policy planning element, which is the proposed action plan updated in step S101, is Plan: [mountain, -]. In this case, the difference is Plan: [mountain, -].
[0229] In step S103, the CPU 11A determines, based on the calculation results of step S102, whether there is a difference between the input policy planning elements, which are the proposed action plan updated in step S101, and the currently set policy planning elements. Since there is a difference, the process proceeds to step S104.
[0230] In step S104, the CPU 11A determines whether interaction is necessary. That is, it determines whether the input policy planning elements, which are the proposed action plan updated in step S103, need to be confirmed with agent AG. Specifically, it refers to the confirmation requirement table 13D and determines whether the input policy planning elements, which are the proposed action plan updated in step S103, are input policy planning elements that need to be confirmed with a person.
[0231] The Confirmation Requirement Table 13D is a table that specifies whether or not confirmation from Agent AG is required when a difference occurs between the input policy plan elements as a proposed action plan and the elements currently set as the action plan.
[0232] Specifically, as shown in Figure 27, the confirmation requirement table 13D is a table that shows the correspondence between the input policy plan elements as updated action plan proposals, the changes from the elements currently set as action plans, the changes, the reasons for the changes, the content of the changes, the magnitude of the changes, and whether confirmation is required.
[0233] For example, in the above example, the change is in "Plan", the reason for the change is "reflecting human instructions", the content of the change is "mountain road", and the magnitude of the change is "-", so confirmation is not required.
[0234] In step S105, based on the result of the interaction requirement determination in step S104, it is determined whether interaction is necessary, that is, whether confirmation from a person is required. In this case, confirmation from a person is not required, so the process proceeds to step S107.
[0235] In step S107, the CPU 11A adopts the input policy plan elements as proposed action plans updated in step S101 and sets them as elements for action plans. Here, [Goal: [time], Action: [route x] are set as elements for action plans.
[0236] In step S108, the CPU 11A determines the inconsistency between layers. Here, it is assumed that an inconsistency has occurred because the likelihood of the elements in the higher layer that are current action plan proposals for the input action plan elements that have been input into the action plan model 13C is below a threshold.
[0237] In step S109, the CPU 11A determines whether or not it determined in step S108 that an inconsistency occurred between layers. Here, it is determined that an inconsistency has occurred, so the process proceeds to step S110.
[0238] In step S110, the CPU 11A inputs the input policy planning elements as action plans updated in step S107 into the action planning model 13C, and sets the element with the highest likelihood of being an action plan proposal at a higher level from the output data output from the action planning model 13C as the input policy planning element for the action plan proposal. Here, [Plan: [mountain, -]] is input into the action planning model 13C, and (Goal: [load], Action: [route x]) output from the action planning model 13C is set as the input policy planning element for the action plan proposal. Note that "load" is an element that means reducing the operating load. In this way, the content of Goal was initially "time", but the content of Goal is changed to "load" by human instruction.
[0239] Next, we will explain a specific example of the case where the process flows in the order of steps S102-S106 and S100-S106 in the flowchart of Figure 3. In the following, we will assume that agent AG is a human. Furthermore, we will assume that the element representing the current action policy of robot RB's Goal is set to "time", and that in step S101 the element representing the action policy of Goal is set to "load".
[0240] In step S102, the CPU 11A determines whether the input policy planning elements, which are the proposed action plan for Goal updated in step S101, deviate from the currently set policy planning elements, which are the action plan for Goal. Here, it is determined that the input policy planning elements, which are the proposed action plan for Goal updated in step S101, deviate from the currently set policy planning elements, which are the action plan for Goal.
[0241] In step S103, the CPU 11A determines that there is a discrepancy based on the judgment result in step S102, and proceeds to step S104.
[0242] In step S104, the CPU 11A determines whether it needs to confirm with agent AG the "load" element, which is an element of the proposed action plan for Goal updated in step S101, by referring to the confirmation necessity table 13D. Here, the change is in "Goal", the reason for the change is "robot's spontaneous change (unconfirmed)", the change content is "-", and the change range is "-", so confirmation is required.
[0243] In step S105, based on the result of the interaction requirement determination in step S104, it is determined whether interaction is necessary, that is, whether confirmation from a person is required. In this case, confirmation from a person is required, so the process proceeds to step S106.
[0244] In step S106, the CPU 11A performs an interaction. Here, it communicates with the person by outputting the voice message "Would you prefer the easy route?" from the robot RB.
[0245] In step S100, the CPU 11A inputs the voice input signal from the person detected by the behavior sensor BS to the trained classification model 13B and obtains the action plan hierarchy and elements corresponding to the voice input signal from among the multiple action plan hierarchys included in the relational network as an action plan hierarchy structure. Here, let's assume the person says "Yes". In this case, the trained classification model 13B outputs "load" as the input policy planning element as the proposed action plan for Goal. This is because the person gave an affirmative response of "Yes" to the question "Would you prefer the easy way?" from the robot control device 10, and therefore "load" is output as the input policy planning element as the proposed action plan for Goal, which is the same as the question.
[0246] In step S101, the CPU 11A updates the input policy planning elements as proposed action plans. Here, the Goal: [load, -] is input to the action planning model 13C, and the output data output by the action planning model 13C with the highest likelihood ([Plan: [city road], Action: [route Y]) is selected as the input policy planning element as a proposed action plan. This input policy planning element as a proposed action plan means that the content of Plan is set to a city road and the content of Action is set to route Y.
[0247] In step S102, the CPU 11A calculates whether the input policy planning elements, which are the proposed action plan updated in step S101, deviate from the elements currently set as action plans, i.e., the difference. Here, the difference is [Plan, city road, -] and [Goal, load, -].
[0248] In step S103, the CPU 11A determines, based on the calculation results of step S102, whether there is a difference between the input policy planning elements, which are the proposed action plan updated in step S101, and the currently set policy planning elements, which are the action plan. Since there is a difference, the process proceeds to step S104.
[0249] In step S104, the CPU 11A determines whether interaction is necessary. That is, it determines whether it is necessary to have a person confirm the input policy plan elements, which are the updated action plan proposal in step S103. Here, the reason for the change in the "Plan" section is "voluntary change (unconfirmed)", the content of the change is "city road", and the change range is "-", so confirmation is required. Also, the reason for the change in the "Goal" section is "voluntary change (confirmed)", the content of the change is "-", and the change range is "-", so confirmation is not required.
[0250] In step S105, based on the result of the interaction requirement determination in step S104, it is determined whether interaction is necessary, that is, whether confirmation from a person is required. In this case, confirmation from Plan is required, so the process proceeds to step S106.
[0251] In step S106, the CPU 11A performs an interaction. Here, it communicates with a person by outputting the voice message "How about going through the city?" from the robot RB.
[0252] Thus, even if a car navigation system sets a route that takes you through mountain roads with the aim of saving time, if the driver chooses to take an easier route with the aim of reducing the burden on the driver, urban roads will be selected as the route according to the driver's intention.
[0253] Conventionally, even if the route set by a car navigation system contradicted a person's intentions, the car navigation system could not understand the person's intentions and correct the route. Therefore, the person had to set the route themselves. In contrast, according to this embodiment, the system can infer and confirm the person's higher-level intentions based on the route selected by the person, and set a route that is consistent with the confirmation result.
[0254] Although each embodiment has been described above, these embodiments are merely illustrative examples of the configurations of the present disclosure. The present disclosure is not limited to the specific forms described above, and various modifications are possible within the scope of its technical concept.
[0255] In addition, the learning and control processes that the CPU reads and executes in each of the above embodiments may be executed by various processors other than the CPU. Examples of such processors include dedicated electrical circuits, which are processors with circuit configurations specifically designed to execute particular processes, such as FPGAs (Field-Programmable Gate Arrays), PLDs (Programmable Logic Devices) whose circuit configurations can be changed after manufacturing, and ASICs (Application Specific Integrated Circuits). Furthermore, the equipment management process may be executed by one of these various processors, or by a combination of two or more processors of the same or different types (for example, multiple FPGAs, and a combination of a CPU and an FPGA). More specifically, the hardware structure of these various processors is an electrical circuit that combines circuit elements such as semiconductor elements.
[0256] The following are additional notes regarding this disclosure.
[0257] (Note 1) A relational network that defines the respective hierarchical relationships between multiple elements, further comprising a specific hierarchy including multiple policy or plan elements relating to the operation of a controlled device, wherein each element has a hierarchical relationship with at least one other element, where it is directly superior or directly inferior, and the hierarchical relationship includes information about the strength of the relationship, wherein for policy or plan elements and operation elements that are directly in a hierarchical relationship, the policy or plan element is positioned higher than the operation element, the policy or plan element is positioned at the top of the chain of elements connected by a hierarchical relationship, and the operation element is positioned at the bottom of the chain of elements connected by a hierarchical relationship, and further comprising a relational network having a specific hierarchy including multiple policy or plan elements that are not in a hierarchical relationship with each other, and comprising one set policy or plan element included in the specific hierarchy, Information processing device comprising: a storage unit that stores a relationship network in which a series of setting elements are pre-configured, each being a setting operation element directly associated with the setting policy planning element according to the strength of the relationship, or a chain of setting elements directly associated with the strength of the relationship, where the lowest element in the chain is a setting operation element; a classification unit that classifies information acquired from the outside into corresponding elements in the relationship network; an extraction process that extracts an input policy planning element that belongs to a specific hierarchy among the elements directly associated with the higher level, or a chain of elements directly associated with the higher level, based on the strength of the relationship, from the classified elements classified by the classification unit; a determination process that determines whether the currently configured setting policy planning element of a specific hierarchy matches the extracted input policy planning element; and if the determination process determines that they do not match, a confirmation output that confirms whether to change the setting to the input policy planning element.(Note 2) The processing unit, in response to the input confirmation result input for the input confirmation output, resets the input policy plan element as the 1 setting policy plan element included in the specific layer, and resets a plurality of elements as the series of setting elements, which are either a 1 operation element directly associated with the input policy plan element according to the strength of the relationship, or a chain of 1 elements directly associated with the input policy plan element according to the strength of the relationship, where the lowest element in the chain is a setting operation element. (Note 3) The relationship network includes a first layer containing a plurality of operation elements for operation content, a second layer containing a plurality of policy plan elements for policies or plans that are higher than the operation content of the first layer, and a third layer containing a plurality of higher-level policy plan elements for higher-level policies or plans that are higher than the policy plans of the second layer, wherein the specific layer is the second layer or the third layer. (Note 4) The classification unit classifies the information obtained from the outside into corresponding hierarchies and elements in the relational network, and the processing unit, instead of the extraction process and the judgment process, receives the hierarchies and elements given by the classification unit, calculates the difference between the strength of the relationship between the input element and the element currently set in a higher hierarchy than the input hierarchy, and the strength of the relationship between the element currently set in the input hierarchy and the element in a higher hierarchy than the currently set hierarchy, determines whether the calculated difference exceeds a threshold, and if it is determined that the threshold is exceeded, outputs a confirmation output to confirm whether to change the setting to the input policy planning element.(Note 5) The information processing device according to any one of Notes 1 to 4, wherein, in the judgment process, it is determined that the elements are consistent, and it is determined that the elements of 1 that are directly associated with the input policy plan element of 1, which is an element belonging to the specific hierarchy, are inconsistent with the elements of 1 that are directly associated with the input policy plan element of 1, which is an element belonging to the specific hierarchy, among the elements that are directly or chain-linked to be associated with the classified elements based on the strength of their relationship.(Appendix 6) A relationship network that defines the respective hierarchical relationships between multiple elements, wherein a relationship network that includes at least multiple operational elements for the operation of a controlled device and multiple policy or plan elements for a policy or plan, wherein each element has a hierarchical relationship with at least one other element, where it is directly superior or directly inferior, and the hierarchical relationship includes information regarding the strength of the relationship, wherein for policy or plan elements and operational elements that are directly in a hierarchical relationship, the policy or plan element is positioned higher than the operational element, the policy or plan element is positioned at the top of the chain of elements connected by a hierarchical relationship, and the operational element is positioned at the bottom of the chain of elements connected by a hierarchical relationship, wherein a relationship network that defines the respective hierarchical relationships between multiple elements, wherein a chain of 1 configuration operational elements that are directly associated with one top-level configuration policy or plan element according to the strength of the relationship, or a chain of 1 configuration elements that are directly associated with one top-level configuration policy or plan element, where the lowest element in the chain is a configuration operational element, is pre-configured, and a series of configuration elements that are linked vertically in a hierarchical relationship according to the strength of the relationship is pre-configured, and a classification unit that classifies information acquired from the outside into the corresponding elements in the relationship network, Information processing device comprising: an extraction process that extracts a series of input corresponding elements that are directly or chained to a higher level based on the strength of their relationship to the classified elements, up to the highest-level policy planning element based on the strength of their relationship; a determination process that determines whether there are any elements that are consistent between the series of input corresponding elements and the series of setting elements; and a processing unit that outputs a confirmation output confirming whether the setting should be changed to one of the input corresponding elements from the series of input corresponding elements if the determination process determines that there are no elements that are consistent.(Addendum 7) The information processing device according to Addendum 6, wherein, if the processing unit determines in the determination process that there is a matching element, it determines whether the lower input corresponding element directly below the matching element among the series of input corresponding elements and the lower setting element directly below the matching element among the series of setting elements are matching, and if the lower input corresponding element and the lower setting element are not matching, it outputs a confirmation output to confirm whether to change the setting to part or all of the series of input corresponding elements. (Addendum 8) The information processing device according to claim 7, wherein, when the lower input corresponding element and the lower setting element are not matching, the processing unit outputs the confirmation output if the mismatched element is the policy planning element, and does not output the confirmation output if the mismatched element is the operation element. (Addendum 9) The information processing device according to Addendum 6 or Addendum 7, wherein the processing unit changes the setting to the series of input corresponding elements in response to the confirmation result input for the input confirmation output. (Note 10) The information processing device according to any one of Note 1 to 9, comprising a proposal suitability calculation unit that calculates the suitability of the proposed input policy planning elements or input operation elements extracted by the processing unit in the extraction process, wherein the processing unit sets the input policy planning elements or input operation elements of the controlled device based on the proposal suitability. (Note 11) The information processing device according to any one of Note 1 to 10, comprising a proposal suitability acquisition unit that obtains a proposal suitability weighted by the recommendation degree of multiple viewpoints for the input policy planning elements or input operation elements extracted by the processing unit in the extraction process, wherein the processing unit sets the input policy planning elements or input operation elements of the controlled device based on the proposal suitability. (Note 12) The information processing device according to any one of Note 1 to 11, comprising an interaction adjustment unit that adjusts at least one of the output timing of the confirmation output and the confirmation output for the input policy planning elements or input operation elements extracted by the processing unit in the extraction process.(Note 13) An information processing device according to any one of Note 1 to 12, comprising an immediate execution determination unit that, when a difference occurs between the input policy planning element or input operation element of the controlled device and the currently set setting policy planning element or setting operation element, immediately determines whether or not to adopt the input policy planning element or input operation element without sending the confirmation output to an external agent regarding the difference. (Note 14) An information processing device according to any one of Note 1 to 13, comprising an agent model that estimates the operation trend of the agent based on the operation history of the agent, wherein the processing unit sets the input policy planning element or input operation element of the controlled device based on the operation trend of the agent estimated by the agent model.(Note 15) An information processing method that includes the following steps: a computer classifies information acquired from an external source into corresponding elements in a relational network; extracts one input policy / plan element that belongs to a specific hierarchy among one element that is directly associated with a higher level based on the strength of the relationship, or a chain of elements that are directly associated with a higher level, from the classified elements; performs a judgment process to determine whether the currently set policy / plan element, which is an element of a specific hierarchy, and the extracted input policy / plan element are consistent; and if the judgment process determines that they are not consistent, outputs a confirmation output to confirm whether to change the setting to the input policy / plan element, wherein the relational network includes at least a plurality of operation elements for the operation of a controlled device and a plurality of policy / plan elements for policies or plans, each element has a hierarchical relationship in which it is directly superior or directly inferior to at least one other element, the hierarchical relationship includes information regarding the strength of the relationship, and for policy / plan elements and operation elements that have a direct hierarchical relationship, the policy / plan element is positioned on the higher side relative to the operation element. An information processing method comprising a relational network having a specific hierarchy that includes multiple policy and plan elements that are not in a hierarchical relationship with one another, wherein a policy and plan element is placed at the top of a chain of elements that are connected in a hierarchical relationship, and an action element is placed at the bottom of a chain of elements that are connected in a hierarchical relationship, and the relational network having a specific hierarchy that includes multiple policy and plan elements that are not in a hierarchical relationship with one another, wherein the relational network is a relational network in which a series of setting elements are pre-configured to correspond in a hierarchical relationship and are linked in a hierarchical relationship according to the strength of the relationship, and the lowest element in the chain is a setting action element.(Note 16) An information processing program that causes a computer to perform the following processes: classify information obtained from an external source into corresponding elements in a relational network; extract one input policy / plan element that belongs to a specific hierarchy among one element that is directly associated with a higher level based on the strength of the relationship, or a chain of one element that is directly associated with a higher level, from the classified elements; perform a judgment process to determine whether the currently set policy / plan element, which is an element of a specific hierarchy, and the extracted input policy / plan element are consistent; and if the judgment process determines that they are not consistent, output a confirmation output to confirm whether to change the setting to the input policy / plan element, wherein the relational network includes at least a plurality of operation elements for the operation of a controlled device and a plurality of policy / plan elements for policies or plans, each element has a hierarchical relationship in which it is directly superior or directly inferior to at least one other element, the hierarchical relationship includes information regarding the strength of the relationship, and for policy / plan elements and operation elements that have a direct hierarchical relationship, the policy / plan element is positioned on the higher side relative to the operation element. An information processing program in which a relational network is configured such that a policy / planning element is placed at the top of a chain of elements connected in a hierarchical relationship, an action element is placed at the bottom of a chain of elements connected in a hierarchical relationship, a relational network that defines the hierarchical relationships between multiple elements, and furthermore, a relational network having a specific hierarchy that includes multiple policy / planning elements that are not in a hierarchical relationship with each other, and a relational network in which a series of setting elements are pre-configured that are linked vertically according to the strength of the relationship, and one setting action element is directly associated with the one setting policy / planning element in a lower position according to the strength of the relationship, or a chain of one setting element that is directly associated with a lower position according to the strength of the relationship, with the lowest element in the chain being a setting action element.(Note 17) An information processing method comprising: a classification unit that classifies information acquired from an external source into corresponding elements in a relational network; an extraction process that extracts a series of input corresponding elements that are directly or chained to a higher level based on the strength of the relationship with the classified elements, up to the highest-level policy / planning element, based on the strength of the relationship; a determination process that determines whether there are any elements that are consistent between the series of input corresponding elements and the series of setting elements; and if the determination process determines that there are no elements that are consistent, an output that confirms whether the setting should be changed to any of the input corresponding elements in the series of input corresponding elements, wherein the relational network includes at least a plurality of operation elements for the operation of a controlled device and a plurality of policy / planning elements for a policy or plan, each element has a hierarchical relationship in which it is directly higher or lower than at least one other element, the hierarchical relationship includes information regarding the strength of the relationship, and for policy / planning elements and operation elements that are directly hierarchical, the policy / planning element is positioned higher than the operation element. An information processing method comprising a relational network in which a series of setting elements are pre-configured to correspond vertically and are linked vertically, with a policy / planning element positioned at the top of a chain of elements connected in a hierarchical relationship, and an action element positioned at the bottom of a chain of elements connected in a hierarchical relationship, and the relational network defining the hierarchical relationships between multiple elements, wherein for one top-level setting policy / planning element, one setting action element is directly associated with a lower level according to the strength of the relationship, or a chain of one setting element directly associated with a lower level according to the strength of the relationship, and the lowest element in the chain is a setting action element, and a series of setting elements are pre-configured to correspond vertically and are linked according to the strength of the relationship.(Note 18) An information processing program that causes a computer to perform the following: a classification unit that classifies information acquired from an external source into corresponding elements in a relational network; an extraction process that extracts the classified elements classified by the classification unit, and one or more elements that are directly or chain-linked to higher levels based on the strength of their relationship to the classified elements, and which are chained based on the strength of their relationship down to the highest-level policy and planning element; a determination process that determines whether there are any elements that are consistent between the series of input and planning elements; and if the determination process determines that there are no elements that are consistent, a confirmation output that confirms whether the setting should be changed to any of the input and planning elements in the series of input and planning elements, wherein the relational network includes at least a plurality of operation elements for the operation of a controlled device, and a plurality of policy and planning elements for a policy or plan, each element has a hierarchical relationship in which it is directly higher or lower than at least one other element, the hierarchical relationship includes information about the strength of the relationship, and for policy and planning elements and operation elements that are directly hierarchical, the policy and planning element is positioned higher than the operation element. An information processing program that is a relational network in which a series of setting elements are pre-configured to correspond vertically and are linked vertically, with a policy / planning element placed at the top of a chain of elements connected in a hierarchical relationship, and an action element placed at the bottom of a chain of elements connected in a hierarchical relationship, and the hierarchical relationships between multiple elements being defined, wherein for one top-level setting policy / planning element, one setting action element is directly associated with a lower level according to the strength of the relationship, or a chain of one setting element directly associated with a lower level according to the strength of the relationship, and the lowest element in the chain is a setting action element.
[0258] Furthermore, the disclosures of Japanese Patent Application No. 2024-200080 and Japanese Patent Application No. 2025-003658 are incorporated herein by reference in their entirety. In addition, all documents, patent applications, and technical standards described herein are incorporated herein by reference to the same extent as if the incorporation of each individual document, patent application, and technical standard were specifically and individually noted.
Claims
1. A relationship network that includes at least a plurality of operational elements for the operation of a controlled device, and a plurality of policy or plan elements for a policy or plan, wherein each element has a hierarchical relationship with at least one other element, where it is directly superior or directly inferior, and the hierarchical relationship includes information about the strength of the relationship, where, for policy or plan elements and operational elements that are directly in a hierarchical relationship, the policy or plan element is positioned higher than the operational element, the policy or plan element is positioned at the top of the chain of elements connected by a hierarchical relationship, and the operational element is positioned at the bottom of the chain of elements connected by a hierarchical relationship, and the relationship network that defines the respective hierarchical relationships between the plurality of elements, further comprising a specific hierarchy including a plurality of policy or plan elements that are not in a hierarchical relationship with each other, wherein the relationship network includes one set policy or plan element included in the specific hierarchy, and a set operational element that is directly associated with the one set policy or plan element in a hierarchical relationship, or a chain of set elements that are directly associated with the one set policy or plan element in a hierarchical relationship, where the lowest element in the chain is the set operational element, and a series of set elements that are linked vertically in a hierarchical relationship Information processing device comprising: a classification unit that classifies information obtained from an external source into corresponding elements in the relational network; an extraction process that extracts one input policy plan element that belongs to a specific hierarchy among one element that can be directly associated with a higher level based on the strength of the relationship, or a chain of one element that can be directly associated with a higher level, from the classified elements classified by the classification unit; a determination process that determines whether the currently set policy plan element, which is an element of a specific hierarchy, and the extracted input policy plan element are consistent; and if the determination process determines that they are not consistent, a processing unit that outputs a confirmation output to confirm whether to change the setting to the input policy plan element.
2. The information processing apparatus according to claim 1, wherein the processing unit, in response to the input confirmation result for the input confirmation output, reconfigures the input policy planning element as the 1 setting policy planning element included in the specific hierarchy, and reconfigures a plurality of elements as the series of setting elements, which are either 1 operation elements directly associated with the input policy planning element according to the strength of their relationship, or 1 chain of elements directly associated with the input policy planning element according to the strength of their relationship, wherein the lowest element in the chain is a setting operation element.
3. The information processing apparatus according to claim 1, wherein the relational network includes a first layer containing multiple operational elements for operational content, a second layer containing multiple policy-planning elements for policies or plans that are higher than the operational content of the first layer, and a third layer containing multiple higher-level policy-planning elements for higher-level policies or plans that are higher than the policies or plans of the second layer, and the specific layer is the second layer or the third layer.
4. The information processing apparatus according to claim 3, wherein the classification unit classifies the information obtained from the outside into corresponding hierarchy and elements in the relational network, and the processing unit, instead of the extraction process and the judgment process, receives the hierarchy and elements given by the classification unit, calculates the difference between the strength of the relationship between the input element and the element currently set in a higher hierarchy than the input hierarchy, and the strength of the relationship between the element currently set in the input hierarchy and the element in a higher hierarchy than the currently set hierarchy, determines whether the calculated difference exceeds a threshold, and if it is determined that the threshold is exceeded, outputs a confirmation output to confirm whether to change the setting to the input policy planning element.
5. The information processing apparatus according to claim 1, in the case where the processing unit determines in the judgment process that the elements are consistent, and it is determined that the pre-set elements of 1 that are directly associated with the setting policy planning elements of 1 included in the specific hierarchy according to the strength of the relationship and the elements that are directly associated with the input policy planning elements of 1, which are elements belonging to the specific hierarchy, are inconsistent, the processing unit changes the setting to the elements that are directly associated with the classified elements, either higher or lower, according to the strength of the relationship.
6. A relationship network that defines the respective hierarchical relationships between multiple elements, comprising at least multiple operational elements for the operation of a controlled device and multiple policy or plan elements for a policy or plan, wherein each element has a hierarchical relationship with at least one other element, where it is directly superior or directly inferior, and the hierarchical relationship includes information regarding the strength of the relationship, wherein for policy or plan elements and operational elements that are directly in a hierarchical relationship, the policy or plan element is positioned higher than the operational element, the policy or plan element is positioned at the top of the chain of elements connected by a hierarchical relationship, and the operational element is positioned at the bottom of the chain of elements connected by a hierarchical relationship, wherein a relationship network that defines the respective hierarchical relationships between multiple elements, wherein for one top-level set policy or plan element, one set operational element is directly associated with a lower position according to the strength of the relationship, or a chain of one set element is directly associated with a lower position according to the strength of the relationship, and the lowest element in the chain is a set operational element, and a series of set elements that are linked vertically and associated according to the strength of the relationship are pre-configured, and a classification unit that classifies information acquired from the outside into the corresponding elements in the relationship network, Information processing device comprising: an extraction process that extracts a series of input corresponding elements that are directly or chained to a higher level based on the strength of their relationship to the classified elements, up to the highest-level policy planning element based on the strength of their relationship; a determination process that determines whether there are any elements that are consistent between the series of input corresponding elements and the series of setting elements; and a processing unit that outputs a confirmation output confirming whether the setting should be changed to one of the input corresponding elements from the series of input corresponding elements if the determination process determines that there are no elements that are consistent.
7. The information processing device according to claim 6, wherein, if the processing unit determines in the determination process that there is a matching element, it determines whether the lower input corresponding element directly below the matching element among the series of input corresponding elements and the lower setting element directly below the matching element among the series of setting elements are matching, and if the lower input corresponding element and the lower setting element are not matching, it outputs a confirmation output to confirm whether to change the setting to part or all of the series of input corresponding elements.
8. The information processing apparatus according to claim 7, wherein, when the lower input corresponding element and the lower setting element are inconsistent, the processing unit outputs the confirmation output if the inconsistent element is the policy planning element, and does not output the confirmation output if the inconsistent element is the operation element.
9. The information processing apparatus according to claim 6, wherein the processing unit changes the settings of the series of input corresponding elements in accordance with the input confirmation result for the input confirmation output.
10. The information processing apparatus according to claim 1, wherein the processing unit comprises a proposal suitability calculation unit that calculates the suitability of the proposed input policy planning elements or input operation elements extracted in the extraction process, and the processing unit sets the input policy planning elements or input operation elements of the controlled device based on the proposal suitability.
11. The information processing apparatus according to claim 1, comprising a proposal suitability acquisition unit which obtains a proposal suitability weighted by the recommendation degree of multiple viewpoints for the input policy planning elements or input operation elements extracted in the extraction process, wherein the processing apparatus sets the input policy planning elements or input operation elements of the controlled device based on the proposal suitability.
12. The information processing apparatus according to claim 1, further comprising an interaction adjustment unit that adjusts at least one of the output timing of the confirmation output and the confirmation output with respect to the input policy planning elements or input operation elements extracted by the processing unit in the extraction process.
13. The information processing apparatus according to claim 1, further comprising an immediate execution determination unit that, when a difference occurs between the input policy planning element or input operation element of the controlled device and the currently set setting policy planning element or setting operation element, immediately determines whether or not to adopt the input policy planning element or input operation element without sending the confirmation output to an external agent regarding the difference.
14. The information processing apparatus according to claim 1, comprising an agent model that estimates the operating trends of the agent based on the agent's operating history, wherein the processing unit sets input policy planning elements or input operation elements of the controlled device based on the operating trends of the agent estimated by the agent model.
15. An information processing method comprising: a computer classifying information acquired from an external source into corresponding elements in a relational network; extracting one input policy / plan element that belongs to a specific hierarchy among one element that is directly associated with a higher level based on the strength of the relationship, or a chain of elements that are directly associated with a higher level, based on the strength of the relationship; performing a judgment process to determine whether the currently set policy / plan element, which is an element of a specific hierarchy, and the extracted input policy / plan element are consistent; and if the judgment process determines that they are not consistent, outputting a confirmation output to confirm whether to change the setting to the input policy / plan element, wherein the relational network includes at least a plurality of operation elements about the operation of a controlled device and a plurality of policy / plan elements about policies or plans, each element has a hierarchical relationship in which it is directly superior or directly inferior to at least one other element, the hierarchical relationship includes information about the strength of the relationship, and for policy / plan elements and operation elements that are directly in a hierarchical relationship, the policy / plan element is positioned higher than the operation element, and the policy / plan element is positioned at the top of the chain of elements connected by a hierarchical relationship. An information processing method comprising a relational network having a specific hierarchy, wherein an operational element is placed at the lowest level of a chain of elements connected in a hierarchical relationship, defining the hierarchical relationships between multiple elements, and further comprising a specific hierarchy including multiple policy and planning elements that are not in a hierarchical relationship with one another, wherein the relational network is a chain of one setting policy and planning element included in the specific hierarchy, and one setting operational element that is directly associated with the one setting policy and planning element in a lower level according to the strength of the relationship, or one setting element that is directly associated with the lower level according to the strength of the relationship, and the lowest element in the chain is a setting operational element, and a series of setting elements that are linked vertically according to the strength of the relationship are pre-configured.
16. An information processing program that causes a computer to perform the following processes: classify information obtained from an external source into corresponding elements in a relational network; extract one input policy / plan element that belongs to a specific hierarchy among one element that is directly associated with a higher level based on the strength of the relationship, or a chain of elements that are directly associated with a higher level, based on the strength of the relationship; perform a judgment process to determine whether the currently set policy / plan element, which is an element of a specific hierarchy, and the extracted input policy / plan element are consistent; and if the judgment process determines that they are not consistent, output a confirmation output to confirm whether to change the setting to the input policy / plan element, wherein the relational network includes at least a plurality of operation elements about the operation of a controlled device and a plurality of policy / plan elements about policies or plans, each element has a hierarchical relationship in which it is directly superior or directly inferior to at least one other element, the hierarchical relationship includes information about the strength of the relationship, and for policy / plan elements and operation elements that are directly in a hierarchical relationship, the policy / plan element is positioned higher than the operation element, and the policy / plan element is positioned at the top of the chain of elements connected by a hierarchical relationship. An information processing program in which an operational element is placed at the lowest level of a chain of elements connected in a hierarchical relationship, is a relational network that defines the hierarchical relationships between multiple elements, and further has a specific hierarchy that includes multiple policy and planning elements that are not in a hierarchical relationship with each other, and is a relational network in which a series of setting elements are pre-configured that are linked vertically according to the strength of the relationship, and one setting operational element is directly associated with the one setting policy and planning element in a lower level according to the strength of the relationship, or one setting element is directly associated with a lower level according to the strength of the relationship, and the lowest element in the chain is a setting operational element.
17. An information processing method comprising: a computer classifying information acquired from an external source into corresponding elements in a relational network; an extraction process that extracts the classified elements classified by the classification unit, and one or more elements that are directly or chain-linked to higher levels based on the strength of their relationships with the classified elements, and which are chained down to the highest-level policy / planning element based on the strength of their relationships; a determination process that determines whether there are any elements that are consistent between the series of input-corresponding elements and the series of setting elements; and, if the determination process determines that there are no elements that are consistent, a confirmation output that confirms whether the setting should be changed to any of the input-corresponding elements in the series of input-corresponding elements, wherein the relational network includes at least a plurality of operation elements for the operation of a controlled device, and a plurality of policy / planning elements for policies or plans, each element has a hierarchical relationship in which it is directly higher or lower than at least one other element, the hierarchical relationship includes information regarding the strength of the relationship, and for policy / planning elements and operation elements that are directly hierarchical, the policy / planning element is positioned higher than the operation element. An information processing method comprising a relational network in which a series of setting elements are pre-configured to correspond vertically and are linked vertically, with a policy / planning element positioned at the top of a chain of elements connected in a hierarchical relationship, and an action element positioned at the bottom of a chain of elements connected in a hierarchical relationship, and the relational network defining the hierarchical relationships between multiple elements, wherein for one top-level setting policy / planning element, one setting action element is directly associated with a lower level according to the strength of the relationship, or a chain of one setting element directly associated with a lower level according to the strength of the relationship, and the lowest element in the chain is a setting action element, and a series of setting elements are pre-configured to correspond vertically and are linked according to the strength of the relationship.
18. An information processing program that causes a computer to perform the following: a classification unit that classifies information acquired from an external source into corresponding elements in a relational network; an extraction process that extracts the classified elements classified by the classification unit, and one or more elements that are directly or chain-linked to higher levels based on the strength of their relationships with the classified elements, and which are chained down to the highest-level policy / planning element based on the strength of their relationships; a determination process that determines whether there are any elements that are consistent between the series of input-corresponding elements and the series of setting elements; and if the determination process determines that there are no elements that are consistent, a confirmation output that confirms whether the setting should be changed to any of the input-corresponding elements in the series of input-corresponding elements, wherein the relational network includes at least a plurality of operation elements for the operation of a controlled device, and a plurality of policy / planning elements for policies or plans, each element has a hierarchical relationship in which it is directly higher or lower than at least one other element, the hierarchical relationship includes information about the strength of the relationship, and for policy / planning elements and operation elements that are directly hierarchical, the policy / planning element is positioned higher than the operation element. An information processing program that is a relational network in which a series of setting elements are pre-configured to correspond vertically and are linked vertically, with a policy / planning element placed at the top of a chain of elements connected in a hierarchical relationship, and an action element placed at the bottom of a chain of elements connected in a hierarchical relationship, and the hierarchical relationships between multiple elements being defined, wherein for one top-level setting policy / planning element, one setting action element is directly associated with a lower level according to the strength of the relationship, or a chain of one setting element directly associated with a lower level according to the strength of the relationship, and the lowest element in the chain is a setting action element.