Learning device, control device, control system, learning method and program

JPWO2024180656A5Pending Publication Date: 2025-10-14
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
JP2025503288
Authority / Receiving Office
JP · JP
Patent Type
Applications
Filing Date
2025-08-01
Publication Date
2025-10-14

AI Technical Summary

Technical Problem

Existing learning devices struggle to determine the success or failure of robot operations in dynamic environments where operation results are uncertain due to adjustable parameters, making it difficult to achieve preset motion goals.

Method used

A learning device that receives motion and control parameters, utilizes a level set function learning unit to evaluate motion goals and a high-level controller learning unit to determine control parameters, enabling the robot to achieve target motions by calculating evaluation values for feasibility and reachability.

Benefits of technology

Enables the robot to effectively achieve target motions by determining achievable control parameters, even when initial parameters are modified, ensuring the robot can reach desired states in uncertain environments.

✦ Generated by Eureka AI based on patent content.
Patent Text Reader

Abstract

In the present invention, a level set function learning unit learns a level set function in which an operation parameter stipulating an operation objective and operation environment of a robot, and a control parameter of the robot are input, and in which an evaluation value relating to the achievability of the operation objective based on the operation parameter and the control parameter is output. A high-level controller learning unit learns a high-level controller that determines, for the robot, a control parameter for realizing the objective operation based on the operation parameter, on the basis of the precision of prediction of the control parameter and the level set function.
Need to check novelty before this filing date? Find Prior Art

Description

Learning device, control device, control system, learning method, and storage medium

[0001] The present application relates to a learning device, a control device, a control system, a learning method, and a storage medium.

[0002] Patent Literature 1 describes an information processing device that generates one or more virtual models of moving machines and virtually operates the moving machines in a virtual environment in which the generated virtual models are placed. When there are multiple types of moving machines specified by parameters, the information processing device performs simulations for each type of moving machine. Furthermore, the information processing device uses detection results from sensors in the virtual environment to operate the moving machine to be learned with arbitrary control content and determines whether a preset operation result is obtained.

[0003] Japanese Patent Application Laid-Open No. 2019-153246

[0004] The information processing device described in Patent Document 1 merely determines whether a machine to be trained can achieve a preset operation result. Depending on the operating environment, it may be desirable to adjust the parameters used during training. Since the operation result for the adjusted parameters cannot be known in advance, it is not possible to determine whether the operation will be successful.

[0005] An object of the present application is to provide a learning device, a control device, a control system, a learning method, and a storage medium that solve the above problems.

[0006] According to a first aspect of the present application, a learning device includes a level set function learning unit that learns a level set function that receives as input operation parameters that define the operation target and operation environment of a robot and control parameters of the robot, and outputs an evaluation value regarding the possibility of achieving the operation target based on the operation parameters and the control parameters, and a high-level controller learning unit that learns a high-level controller that determines control parameters for realizing a target operation based on the operation parameters of the robot, based on the level set function and the prediction accuracy of the control parameters.

[0007] According to a second aspect of the present application, a learning method includes a level set function learning step in which a learning device learns a level set function that receives as input operation parameters that define the robot's operation target and operation environment and control parameters of the robot, and outputs an evaluation value regarding the possibility of achieving the operation target based on the operation parameters and the control parameters, and a high-level controller learning step in which a high-level controller that determines control parameters for realizing the robot's target operation based on the operation parameters is learned based on the level set function and the prediction accuracy of the control parameters.

[0008] In a third aspect of the present application, a control device includes a high-level control unit that determines control parameters for realizing a target motion based on motion parameters that define a motion target and an operating environment of the robot, and a motion planning unit that calculates an evaluation value that indicates the feasibility of the target motion based on the motion parameters and the control parameters using a level set function, and when it is determined that the target motion is feasible based on the evaluation value, uses the control parameters for motion control of the robot, and when it is determined that the target motion is not feasible based on the evaluation value, searches for control parameters that make the target motion feasible based on the evaluation value.

[0009] According to one aspect of the present application, even if the control parameters used for learning the high-level controller are modified, it is possible to determine whether or not the target state can be reached.

[0010] 1 is a diagram illustrating an example of the configuration of a control system according to the first embodiment. FIG. 2 is a diagram illustrating an example of known task parameters. FIG. 3 is a diagram illustrating an example of unknown task parameters. FIG. 4 is a diagram illustrating an example of the hardware configuration of a learning device according to the first embodiment. FIG. 5 is a diagram illustrating an example of the hardware configuration of a robot controller according to the first embodiment. FIG. 6 is a diagram illustrating a robot according to the first embodiment. FIG. 7 is a diagram illustrating a system state expressed in an abstract space. FIG. 8 is a diagram illustrating an example of the configuration of a control system related to skill execution according to the first embodiment. FIG. 9 is a diagram illustrating an example of the functional configuration of a learning device related to updating of a skill database according to the first embodiment. FIG. 10 is a diagram illustrating an example of the configuration of a skill learning unit according to the first embodiment. FIG. 11 is a data flow diagram illustrating a data flow in the skill learning unit according to the first embodiment. FIG. 12 is a flowchart illustrating a learning process according to the first embodiment. FIG. 13 is a flowchart illustrating an operation plan according to the first embodiment. FIG. 14 is a data flow diagram illustrating a data flow in the skill learning unit according to the second embodiment. FIG. 15 is a flowchart illustrating a learning process according to the second embodiment. FIG. 16 is a data flow diagram illustrating a data flow in the skill learning unit according to the third embodiment. FIG. 17 is a schematic block diagram illustrating an example of the functional configuration of a system model learning unit according to the third embodiment. FIG. 18 is a flowchart illustrating a learning process according to the third embodiment. FIG. 19 is a schematic block diagram illustrating an example of the minimum configuration of a learning device according to an embodiment of the present application. FIG. 19 is a schematic block diagram illustrating an example of the minimum configuration of a control device according to an embodiment of the present application.

[0011] Hereinafter, embodiments of the present application will be described with reference to the drawings. The following description is not intended to limit the scope of the claims. Furthermore, not all combinations of technical features described in the embodiments are necessarily essential to solving the problem. In other words, even if some combinations are omitted, other combinations may lead to the solution of the problem. For the sake of convenience, in this application, a character consisting of an arbitrary letter "A" with an arbitrary symbol "x" above it will be referred to as "A". x For example, the notation y^ may refer to a character formed by combining the symbol ^ directly above the letter y.

[0012] <First Embodiment> (1) System Configuration An example of the system configuration of a control system 100 according to the first embodiment will be described. FIG. 1 is a diagram showing an example of the configuration of the control system 100 according to the first embodiment. The control system 100 includes a learning device 1, a storage device 2, a robot controller 3, a measuring device 4, and a robot 5. The learning device 1 is connected to the storage device 2 wirelessly or via a cable so as to be able to input and output various types of data. The robot controller 3 is connected to each of the storage device 2, the measuring device 4, and the robot 5 wirelessly or via a cable so as to be able to input and output various types of data. The connection between the learning device 1 and the storage device 2, and the connection between the robot controller 3 and each of the storage device 2, the measuring device 4, and the robot 5 may be direct or may be via a communication network.

[0013] The learning device 1 learns the motion of the robot 5 to execute a given task. In learning the motion, a machine learning method such as self-supervised learning (SSL) is used. The learning device 1 also learns a set of system states that enable the robot 5 to execute the motion to be learned.

[0014] In the present application, the motions and system states to be learned by the learning device 1 are not limited. The learning device 1 can control motions and system information that are controllable and enable learning of that control. The process to be controlled is not limited to motions that involve changes in position or form. For example, the motion of the robot 5 to be controlled may include obtaining measurement data using a sensor.

[0015] The system state refers to the state of the system to be controlled, including not only the robot 5 but also the operating environment of the robot 5. The robot 5 and the operating environment of the robot 5 are sometimes collectively referred to as the target system or simply the system. For a task that involves handling an object, such as a task to grasp an object, the object may also be included in the target system.

[0016] The state of the target system is sometimes called the system state, or simply the state. The system state at the completion of a task as defined in the task is sometimes called the goal state of that task, or simply the goal state. A set of goal states is sometimes called the goal state set. Reaching the goal state of a task is sometimes called accomplishing the task or succeeding in the task. When a task is accomplished by executing a skill, the state at the end of skill execution corresponds to the goal state. The system state at the start of a task is sometimes called the initial state of that task.

[0017] The learning device 1 performs learning on the skills of the robot 5. Each skill is formed by modularizing one or more specific actions of the robot 5. In this application, a task that can be achieved by executing one skill at a time is mainly assumed, and an example will be described in which the learning device 1 learns the skills to achieve that task.

[0018] However, the robot controller 3 may be capable of executing a task that is formed by combining a plurality of skills. For example, the robot controller 3 may plan the execution of a task that is formed by a plurality of subtasks by combining skills for executing the individual subtasks.

[0019] In learning a skill, the learning device 1 may learn a set of states that enables the skill to be executed. The learning device 1 stores information about the skill obtained through learning in the storage device 2, forming a skill database. Information registered in the skill database is sometimes called a skill tuple. A skill tuple is modularized to include various information required to execute the operations that constitute the skill. The learning device 1 generates a skill tuple based on detailed system model information, low-level controller information, and target parameter information stored in the storage device 2.

[0020] The storage device 2 stores various types of information that can be referenced by the learning device 1 and the robot controller 3. The storage device 2 stores, for example, detailed system model information, low-level controller information, target parameter information, and a skill database. Note that the storage device 2 does not necessarily have to be configured separately from the learning device 1 or the robot controller 3. The storage device 2 may be built into the learning device 1 or the robot controller 3. The storage device 2 may be configured to include an external storage device such as a hard disk directly connected to or built into the learning device 1 or the robot controller 3, or a storage medium such as a flash memory. The storage device 2 may be a server device that can execute data communication with the learning device 1 and the robot controller 3. The storage device 2 may also be configured to include multiple storage media, with each storage medium being distributed throughout the control system 100.

[0021] The detailed system model information is information that represents a model of the target system in real space. In this application, the model of the target system in real space may also be referred to as a "detailed system model." The detailed system model may also be referred to as a "detailed" system model to distinguish it from a more abstract "abstract" system model. The detailed system model information may be expressed using a differential equation or difference equation that represents the detailed system model. The detailed system model information may be configured as a simulator program that simulates the operation of the robot 5.

[0022] The low-level controller information is information related to the low-level controller. The low-level controller generates a control input for controlling the actual movement of the robot 5 based on parameter values ​​output from the high-level controller. For example, when the high-level controller generates a trajectory for the robot 5, the low-level controller generates a control input that follows the movement of the robot 5 according to the trajectory. Furthermore, the low-level controller may control the movement of the robot 5 by servo control using PID (Proportional Integral Differential) based on the parameters output from the high-level controller.

[0023] Target parameter information is provided for each skill to be learned by the learning device 1. The target parameter information includes, for example, initial state information, target state / known task parameter information, unknown task parameter information, execution time information, and general constraint information. Here, the variable parts of a task are sometimes referred to as task parameters.

[0024] Task parameters that are expressed as numerical values ​​are sometimes referred to as known task parameters. Examples of known task parameters include the size of an object in a task, such as the size of an object to be grasped when the task is to grasp an object, and the trajectory of the robot 5 for executing the task. Known task parameters can also be treated as skill parameters.

[0025] 2 is a diagram showing examples of known task parameters. Fig. 2 shows an example in which the robot 5 executes a task of grasping a cylindrical object. In this case, the radius and height of the cylindrical object correspond to examples of known task parameters.

[0026] On the other hand, task parameters that are difficult to express numerically are sometimes referred to as unknown task parameters. Examples of unknown task parameters include the shape of an object in a task, such as the shape of an object to be grasped when the task is to grasp an object, and the type of movement of the robot 5 to execute the task, such as the skills required to execute the task.

[0027] 3 is a diagram showing examples of unknown task parameters. Fig. 3 shows an example in which the robot 5 executes a task of grasping objects of various shapes. In this case, the shapes of the objects correspond to an example of unknown parameters.

[0028] In addition, in the present application, it is assumed that the control system 100 handles quantified system states, and the target state is expressed as a numerical value. For example, in the case of a task in which the robot 5 performs pick and place, the target state is expressed as the coordinates of the operation target object being within a predetermined range.

[0029] Initial state information is information that indicates a set of states in which a target skill can be executed. The state at the start of skill execution is also referred to as the initial state of that skill, or simply as the initial state. A set of initial states is also referred to as an initial state set. Parameters of the initial state are also referred to as initial state parameters. For example, if the task of the robot 5 is pick-and-place, the initial hand position and posture correspond to the initial state parameters. The initial hand position corresponds to the position of the end effector of the robot 5 when executing the operation. The initial posture corresponds to the overall shape of the robot 5 before executing the operation. If the robot 5 is configured to include multiple joints, the angle between each pair of two joints that are connected to each other is an element of the initial posture. In this application, the initial state or initial state parameter is defined as x s or x si Here, "i" is a positive integer that represents an identification number that identifies each initial state. Also, the time when the skill execution starts is set to 0, and the initial state is set to x. 0 It may be expressed as:

[0030] The goal state / known task parameter information is information indicating a set of combinations of possible values ​​of a goal state, which is a state that can be reached by executing a target skill, and possible values ​​of known task parameters, which are treated as dynamic parameters of the target skill. For example, in a skill in which the robot 5 grasps an object, a target grasping position and posture may be included as an element of the goal state. The target grasping position refers to the position of the end effector grasping the object to be grasped and may be expressed as a relative target value based on the position of the object. The target posture refers to the target value of the shape of the robot 5 at the start of the grasping operation of the object. The goal state may include information regarding stable grasping conditions such as form closure and force closure as possible values. In this application, a combination of a goal state and a known task parameter value is referred to as a goal state / known task parameter value or a goal state parameter, and β g or β giHere, "i" is a positive integer representing an identification number that identifies each target state / known task parameter value. When the task of the robot 5 is pick-and-place, the final target position of the end effector of the robot 5 and the position of the manipulated object correspond to the target state parameters. In this case, the size of the manipulated object corresponds to the known task parameter.

[0031] By treating the differences in the goal states of tasks and the differences in the known task parameter values ​​as skill parameters, tasks with different goal states and / or different known task parameter values ​​can be executed as a single skill.

[0032] For example, when the learning device 1 performs processing related to skill learning using a predictor, a target state and known task parameter values ​​can be input to the predictor to obtain an output value corresponding to the target state and the known task parameter values. The predictor is configured using a learning model (a model in machine learning) such as a neural network or a Gaussian process (GP).

[0033] Note that, depending on the skill, there may be cases where known task parameters are not set. In this case, the goal state / known task parameter information may be configured as a set of values ​​that the goal state can take. Also, the goal state / known task parameter value β g may be a value indicating the target state.

[0034] The unknown task parameter information is information about unknown task parameters. For example, the unknown task parameter information may indicate a probability distribution of data related to the unknown parameter. If one skill has multiple unknown task parameters, the unknown task parameter information may indicate information about each unknown task parameter. Note that values ​​corresponding to the target state / known task parameters and unknown task parameters may be fixed values ​​or may be variable.

[0035] In this application, the unknown task parameter values ​​are expressed as τ or τ jHere, "j" is a positive integer representing an identification number that identifies the unknown task parameter value. Note that it is difficult to express the unknown task parameter numerically because it is difficult to systematically quantify its value, but it is possible to determine whether the unknown task parameter values ​​are equal. For example, if the unknown task parameter represents the shape of an object, it is possible to determine whether the unknown task parameter values ​​are equal by comparing the shapes of two objects. The control system 100 treats two tasks as the same task if the unknown task parameter values ​​in the two tasks are the same, and treats the two tasks as separate tasks if the unknown task parameter values ​​are different. τ or τ j The above "j" can also be considered as a positive integer representing an identification number that identifies each individual task.

[0036] The execution time information is information regarding the time limit for skill execution. For example, the execution time information may indicate the skill execution time (the time required to execute the skill), or the allowable condition value for the time from the start to the end of skill execution, or both. The general constraint information is information indicating general constraint conditions, such as conditions regarding limits on the range of movement of the robot 5, speed limits, and input limits.

[0037] The skill database is a database that has a skill tuple set for each skill. The skill tuple may include information about a high-level controller for executing the target skill, information about a low-level controller for executing the target skill, and information about a set of states (e.g., initial states for the skill) in which the target skill can be executed and combinations of target states / known task parameter values. The set of states in which the target skill can be executed and target states / known task parameter values ​​is also referred to as an executable state set.

[0038] The feasible state set may be defined in an abstract space obtained by abstracting an actual space. The feasible state set can be expressed using, for example, a level set function estimated using Gaussian Process Regression (GPR) or Level Set Estimation (LSE), or an approximation function of the level set function. In other words, whether the feasible state set includes a combination of a state and a target state / known task parameter value can be determined by whether the value (e.g., the mean value) of the Gaussian process regression for the combination of the state and the target state / known task parameter value, or the value of the approximation function for the combination of the state and the target state / known task parameter value, satisfies a constraint for determining feasibility. The following description will be given using, as an example, a case where a level set function is used as a function representing the feasible state set, but is not limited to this.

[0039] The robot controller 3 is a control device that formulates an operation plan for the robot 5 based on the measurement signals supplied from the measurement device 4, a skill database, etc. The robot controller 3 generates control commands (control inputs) for causing the robot 5 to execute the planned operation, and supplies the control commands to the robot 5.

[0040] For example, the robot controller 3 converts tasks to be executed by the robot 5 into a sequence indicating tasks that the robot 5 can accept for each predetermined time step (time interval). The robot controller 3 then controls the robot 5 based on control commands corresponding to execution commands for the tasks indicated in the generated sequence. The control commands correspond to control inputs output by low-level controllers.

[0041] The measurement device 4 includes one or more sensors and detects the state of the workspace in which the robot 5 executes a task. The sensor may be, for example, a camera, a range sensor, a sonar, or a combination of these. The measurement device 4 supplies the generated measurement signal to the robot controller 3. The measurement device 4 may include a self-propelled or flying sensor (including a drone) that moves within the workspace. The measurement device 4 may also include sensors provided on the robot 5 itself and sensors provided on other objects in the workspace. The measurement device 4 may also include a sensor that detects sound within the workspace, i.e., a microphone. Thus, the measurement device 4 may include various sensors that detect the state of the workspace and are provided at any location.

[0042] The robot 5 performs operations related to the instructed task based on control commands supplied from the robot controller 3. The robot 5 may be a robot that operates in various factories such as an assembly factory or a food factory, or in a logistics site. The robot 5 may be a vertical articulated robot, a horizontal articulated robot, or a robot having another structure. The robot 5 may supply a status signal indicating the status of the robot 5 itself to the robot controller 3. This status signal may be an output signal from a sensor that detects the status (e.g., position, angle, etc.) of the entire robot 5 or a specific part (e.g., joint), or may be a signal indicating the progress of the operation of the robot 5.

[0043] Note that various modifications may be made to the configuration of the control system 100 exemplified in Fig. 1. For example, the robot controller 3 and the robot 5 may be configured as an integrated unit. In another example, at least two of the learning device 1, the storage device 2, and the robot controller 3 may be configured as an integrated unit. Furthermore, the control target of the control system 100 is not limited to the robot 5. Various control targets that the learning device 1 can learn may also be the control targets of the control system 100.

[0044] (2) Hardware Configuration Next, an example of the hardware configuration of the learning device 1 according to this embodiment will be described. Fig. 4 is a diagram showing an example of the hardware configuration of the learning device 1 according to this embodiment. The learning device 1 includes, as hardware, a processor 11, a memory 12, and an interface 13. The processor 11, the memory 12, and the interface 13 are connected via a data bus 10 so as to be able to input and output various types of data.

[0045] The processor 11 functions as a controller (arithmetic unit) that controls the entire learning device 1 by executing a program stored in the memory 12. In this application, "executing a program" or "executing a program" may refer to executing processing instructed by various commands written in the program. The processor 11 is, for example, a processor such as a CPU (Central Processing Unit), a GPU (Graphics Processing Unit), or a TPU (Tensor Processing Unit). The processor 11 is not limited to one processor, and may be configured to include multiple processors. The processor 11 constitutes the computer of the learning device 1.

[0046] The memory 12 includes a storage medium for storing various types of information referenced by the learning device 1 or other hardware. The memory 12 includes various types of volatile and non-volatile memory, such as random access memory (RAM), read-only memory (ROM), and flash memory. The memory 12 stores programs that contain instructions for instructing the processing to be executed by the processor 11. Some of the information stored in the memory 12 may be stored in one or more external storage devices (e.g., storage device 2) that can communicate with the learning device 1, or may be stored in a storage medium that is detachable from other components of the learning device 1.

[0047] The interface 13 includes an interface for connecting the learning device 1 to other devices so that various data can be input and output. The interface 13 may be a wireless interface such as a network adapter for wirelessly transmitting and receiving data to and from other devices, or a hardware interface for wiredly transmitting and receiving data to and from other devices. For example, the interface 13 may be connected to an input device that accepts user input (external input), such as a touch panel, button, keyboard, or voice input device, a display device such as a display or projector, or a sound output device such as a speaker.

[0048] The hardware configuration of the study device 1 is not limited to the configuration illustrated in Figure 4. For example, the study device 1 may incorporate at least one of a display device, an input device, and a sound output device. The study device 1 may also include a storage device 2.

[0049] 5 is a diagram showing an example of the hardware configuration of the robot controller 3 according to this embodiment. The robot controller 3 includes, as hardware, a processor 31, a memory 32, and an interface 33. The processor 31, the memory 32, and the interface 33 are connected via a data bus 30 so as to be able to input and output various types of data.

[0050] The processor 31 executes a program stored in the memory 32 to function as a controller (arithmetic unit) that performs overall control of the robot controller 3. The processor 31 is, for example, a CPU, a GPU, a TPU, or the like. The processor 31 is not limited to a single processor, and may be configured to include multiple processors.

[0051] The memory 32 is configured to include one or more storage media. The memory 32 includes various types of volatile and non-volatile memory, such as RAM, ROM, and flash memory. The memory 32 also stores programs executed by the processor 31. Note that some of the information stored in the memory 32 may be stored in one or more external storage devices (e.g., the storage device 2) that can communicate with the robot controller 3, or may be stored in a storage medium that is detachable from another part of the robot controller 3.

[0052] The interface 33 is an interface for connecting the robot controller 3 to other devices so as to input and output various types of data. The interface 33 may include a wireless interface such as a network adapter for wirelessly transmitting and receiving data to and from other devices, or may include a hardware interface for wiredly transmitting and receiving data to and from other devices.

[0053] The hardware configuration of the robot controller 3 is not limited to the configuration exemplified in Fig. 5. For example, the robot controller 3 may incorporate at least one of a display device, an input device, and a sound output device. The robot controller 3 may also include the storage device 2.

[0054] (3) Abstract Space Next, the abstract space according to this embodiment will be described. The abstract space is a virtual space separate from the real space, and is used by the robot controller 3 to formulate an operation plan for the robot 5 based on the skill tuples.

[0055] Fig. 6 illustrates a robot (manipulator) 5 that grasps an object in real space, and an object to be grasped 6. Fig. 7 expresses the system state illustrated in Fig. 6 in an abstract space.

[0056] Generally, formulating a motion plan for a robot 5 performing a pick-and-place task requires rigorous calculations that take into account the shape of the end effector of the robot 5, the geometric shape of the object to be grasped 6, the grasping position and posture of the robot 5, and the object characteristics of the object to be grasped 6. In contrast, in this embodiment, the robot controller 3 formulates a motion plan in an abstract space in which the states of each object, such as the robot 5 and the object to be grasped 6, are abstractly (simply) represented. In the example of FIG. 7 , the abstract space defines an abstract model 5x corresponding to the end effector of the robot 5, an abstract model 6x corresponding to the object to be grasped 6, and an executable region (see dashed-line frame 60) for the robot 5 to grasp the object to be grasped 6. Note that, as described above, in the abstract space as well, the executable state set is represented as a set of combinations of initial states and target states / known task parameter values ​​in which a skill can be executed. In the example of Figure 7, a set of combinations of initial states and target states / known task parameter values ​​that allow the grasping skill to be executed is illustrated as a grasp operation executable region in a dashed frame 60. The executable state set or target states / known task parameter values ​​correspond to operation parameters that define the robot's operation goals and operating environment. In this way, the state of the end effector and other elements are abstractly expressed by the state of the robot in abstract space. Furthermore, the state of each object corresponding to the operation target or environmental object can also be abstractly expressed in a coordinate system based on the position of a reference object, such as a workbench.

[0057] In this embodiment, the robot controller 3 uses skills to formulate a motion plan in an abstract space that abstracts an actual system. This makes it possible to effectively reduce the computational cost required for motion planning even in multi-stage tasks. In the example of Figure 7, the robot controller 3 formulates a motion plan to execute skills for performing grasping in a graspable area (dashed line frame 60) defined in the abstract space, and generates control commands for the robot 5 based on the formulated motion plan.

[0058] Hereinafter, the state of a system in real space (sometimes referred to herein as a "real system") may be represented as "x" and the state of a system in abstract space (sometimes referred to herein as an "abstract system") may be represented as "x'" to distinguish between them. The state x' is expressed as a vector (sometimes referred to herein as an "abstract state vector"). For example, for a task such as pick-and-place, the abstract state vector includes a vector representing the state of the object to be manipulated (e.g., position, orientation, velocity, etc.), a vector representing the state of the end effector of the manipulable robot 5, and a vector representing the state of environmental objects. The state x' is defined as a state vector that abstractly represents the state of some elements in the real system. Similarly, the target state / known task parameter value in real space is expressed as "β g ”, and the goal state / known task parameter value in the abstract space is defined as “β g These are sometimes distinguished by writing ".

[0059] (4) Control System for Skill Execution Next, a configuration example of a control system for skill execution according to this embodiment will be described. FIG. 8 is a diagram showing a configuration example of a control system for skill execution according to this embodiment. The processor 31 of the robot controller 3 functionally comprises a motion planning unit 34, a high-level control unit 35, and a low-level control unit 36. The system 50 corresponds to an actual system (an actual system including the robot 5). In this application, the high-level control unit 35 is also referred to as a high-level controller, and π H The high-level control unit 35 is an example of a control means. The low-level control unit 36 ​​is also called a low-level controller, and π L The robot controller 3 is an example of a control device that controls the robot 5.

[0060] The motion planning unit 34 formulates a motion plan for the robot 5 based on the state x' in the abstract system and the skill database. The motion planning unit 34 expresses the target state using, for example, a logical expression based on temporal logic. The motion planning unit 34 may express the logical expression using a preset temporal logic such as linear temporal logic, MTL (Metric Temporal Logic), or STL (Signal Temporal Logic). The motion planning unit 34 converts the generated logical expression into a sequence (motion sequence) for each time step. This motion sequence includes, for example, information about the skills used in each time step.

[0061] The high-level control unit 35 recognizes the skill to be executed for each time step from the action sequence generated by the action planning unit 34. Then, the high-level control unit 35 determines the high-level controller "π H ” to generate a parameter “α” to be input to the low-level control unit 36 ​​.

[0062] The high-level control unit 35 calculates the state "x" in the abstract space at the start of the execution of the skill to be executed. s ' and the target state / known task parameter value 'β g The combination of "" is the executable state set "χ 0 If the parameter α belongs to the category ', the control parameter α is generated as shown in equation (1).

[0063]

[0064] The initial state at the start of skill execution is expressed, for example, by a system state in an abstract space. An approximation function g^ of the level set function is set in advance in the robot controller 3. The robot controller 3 determines the initial state x^ by using the approximation function g^ of the level set function depending on whether or not Equation (2) is satisfied. s ' is the feasible state set χ 0 It is possible to determine whether the variable belongs to the '.

[0065]

[0066] Equation (2) can be considered as a constraint that determines the feasibility of a skill from a certain state. Alternatively, the approximation function g is s The approximation function g^ can also be regarded as a model for evaluating whether a goal state can be reached under known task parameter values, based on the approximation function g^. The approximation function g^ is obtained by the learning device 1 performing learning, as will be described later. In this application, the approximation function g^ itself may also be referred to as a level set function.

[0067] The robot controller 3 executes the skill using the low-level control unit 36, and as shown in equation (3), the abstract system state x'(T) at the time when T time has elapsed since the start of execution is converted into the target state set χ' d It is possible to determine whether the set of goal states belongs to χ'. d is a set of goal states in the abstract space after the execution of the skill to be judged. T denotes the execution time of the skill to be judged.

[0068]

[0069] The low-level control unit 36 ​​calculates the control parameter α generated by the high-level control unit 35, the state x of the current real system obtained from the system 50, and the target state / known task parameter value β g The low-level control unit 36 ​​generates the input "u" based on the low-level controller "π L ", the input u is generated as a control command as shown in equation (4).

[0070]

[0071] In addition, the low-level controller π L is not necessarily limited to the input / output relationship that can be expressed by equation (4), and may have a different input / output relationship.

[0072] The low-level control unit 36 ​​recognizes the state x of the robot 5 and the environment, which is recognized using a predetermined state recognition technique based on the measurement signals output from the measurement device 4 (which may include signals acquired from the robot 5). The model of the system 50 calculates the time change x of the state x based on the input u to the robot 5 and the state x. ・ The state equation (5) expresses the function "f" as the output. ・ " represents the differentiation with respect to time or the difference with respect to time.

[0073]

[0074] (5) Updating the Skill Database Next, updating of the skill database according to this embodiment will be described. Fig. 9 is a diagram showing an example of the functional configuration of the learning device 1 related to updating of the skill database according to this embodiment. Functionally, the processor 11 of the learning device 1 includes an abstract system model setting unit 14, a skill learning unit 15, and a skill tuple generation unit 16. Note that Fig. 9 shows an example of the types of data exchanged for each block, but is not limited to this.

[0075] The abstract system model setting unit 14 sets an abstract system model based on the detailed system model information. The set abstract system model is a model obtained by simplifying the detailed system model specified by the detailed system model information. The detailed system model is a model representing the system 50 in FIG. 8.

[0076] The abstract system model is a model that provides an abstract system state indicated by an abstract state vector x' that is constructed based on a detailed system state x in the detailed system model. The motion planning unit 34 uses the abstract system model to formulate a motion plan for the robot 5. The abstract system model setting unit 14 derives the abstract system model from the detailed system model based on, for example, an algorithm stored in advance in the storage device 2 or the like.

[0077] Note that information about the abstract system model may be stored in advance in the storage device 2 or the like. In this case, the abstract system model setting unit 14 may acquire information about the abstract system model from the storage device 2 or the like. The abstract system model setting unit 14 supplies information about the set abstract system model to the skill learning unit 15 and the skill tuple generation unit 16. In the present application, the abstract system model information and the detailed system model information may be collectively referred to as system model information.

[0078] The skill learning unit 15 learns control related to skill execution based on the abstract system model set by the abstract system model setting unit 14, and the detailed system model information, low-level controller information, and target parameter information stored in the storage device 2. The skill learning unit 15 performs learning of control related to skill execution based on, for example, the high-level controller π H is output from the low-level controller π L The skill learning unit 15 learns the control parameter α to be provided to the training data. As will be described later, the skill learning unit 15 also learns the level set function and acquires training data for learning the control parameter α. At this time, the skill learning unit 15 evaluates, for example, the prediction accuracy of the level set function.

[0079] The skill tuple generation unit 16 generates an executable state set χ obtained by learning in the skill learning unit 15. 0 ' and the high-level controller π H The skill tuple generation unit 16 generates a set (tuple) of associated information as a skill tuple, which includes information about the abstract system model set by the abstract system model setting unit 14, information about the low-level controller information, and target parameter information. The skill tuple generation unit 16 then registers the generated skill tuple in a skill database. The data in the skill database is used by the robot controller 3 to control the robot 5.

[0080] The functions of the abstract system model setting unit 14, the skill learning unit 15, and the skill tuple generation unit 16 can be realized, for example, by the processor 11 executing a predetermined program. Alternatively, the program to be executed may be pre-recorded on a non-volatile storage medium, and the processor 11 may install and execute the program to realize these functions. Note that some or all of these functions are not limited to being realized by software execution, but may also be realized by hardware or a combination of hardware and software. Furthermore, some or all of these functions may be realized using user-programmable integrated circuits, such as field-programmable gate arrays (FPGAs) or microcontrollers. In such cases, these integrated circuits execute programs to realize the above functions. Furthermore, the hardware used to realize the above functions may be other types of hardware, such as application-specific standard producers (ASSPs), application-specific integrated circuits (ASICs), or quantum computer control chips. This also applies to other embodiments described below. Furthermore, each of these functions may be realized by multiple computers working together, for example, using cloud computing technology.

[0081] (6) Configuration of the Skill Learning Unit Next, an example configuration of the skill learning unit 15 according to this embodiment will be described. Fig. 10 is a diagram showing an example configuration of the skill learning unit 15 according to this embodiment. Functionally, the skill learning unit 15 includes a search point set setting unit 210, a data acquisition unit 220, and a learning setting unit 230.

[0082] The search point set setting unit 210 includes a search point set initialization unit 211 and a search point selection unit 212. The data acquisition unit 220 includes a system model setting unit 221, a problem setting calculation unit 222, and a data update unit 223. The learning setting unit 230 includes a level set function learning unit 231, a prediction accuracy evaluation function setting unit 232, a prediction accuracy evaluation unit 233, a controller learning evaluation function setting unit 234, and a high-level controller learning unit 240.

[0083] As described above, the skill learning unit 15 generates training data and uses the generated training data to control the high-level controller π H The skill learning unit 15 also learns the level set function. The search point set setting unit 210 learns the high-level controller π H As a candidate for the task setting to be learned, the initial state x s and the target state / known task parameter value β g The search point set setting unit 210 acquires a plurality of sets of search points each consisting of a combination of the above. The search point set setting unit 210 selects a search point from which training data is to be acquired from among the acquired plurality of candidates. The training data is used for learning how to control the robot 5 by the robot controller 3. The search point set setting unit 210 is an example of a search point setting means.

[0084] The search point set initialization unit 211 sets a search point set. The search point set is a set that includes a plurality of search points as elements. Each search point is initialized by a high-level controller π H The level set function is a candidate for the task setting. More specifically, each search point is a candidate for the initial state x s and the target state / known task parameter value β g The search point set initialization unit 211 sets a predetermined number N (N is a predetermined integer equal to or greater than 2) of search points so that the search points are randomly distributed within a predetermined value range for each task, for example.

[0085] In this application, the search point set is defined as Ξ check Also, each search point is expressed as (x s , βg ) or ξ. s , β g By defining the search point (x), the task setting is specified and the behavior of the robot 5 is determined. s , β g ) can also be considered as a parameter indicating the behavior of the robot 5 for each task.

[0086] The search point selection unit 212 selects a search point set Ξ check The search point ξ to obtain the next training data from i (i is an integer between 1 and N). The search point selection unit 212 selects the selected search point ξ i to the data acquisition unit 220 and the learning setting unit 230. The search point selection unit 212 outputs, for example, the search point set Ξ check Randomly select one search point ξ from the multiple search points that make up i After the level set function g^ has been learned at least once in the learning setting unit 230, the search point selection unit 212 selects the prediction accuracy evaluation function J g^ (ξ i , π h (ξ i )) to maximize the evaluation value calculated using the search point ξ i may be selected.

[0087] Prediction accuracy evaluation function J g^ (ξ i , π h (ξ i )) is a function that indicates the estimation accuracy of the evaluation value obtained for a search point using the level set function g^. The larger the evaluation value of the prediction accuracy evaluation function, the higher the estimation accuracy. The level set function g^ determines an evaluation value that indicates the possibility of reaching the target state. The smaller the evaluation value of the level set function g^, the higher the possibility of reaching the target state. In this case, the search point that gives the evaluation value of the level set function g^ is determined to be able to reach the target state.

[0088] The data acquisition unit 220 receives the search point ξ input from the search point set setting unit 210. i Using the high-level controller π HThe system model setting unit 221 acquires training data for learning the search point ξ and also acquires training data for learning the level set function g. i The system model setting unit 221 sets up a system model for solving an optimal control problem based on the above. The system model setting unit 221 sets up a system model based on the previously set system model information, constraint conditions, low-level controller, target time, solution search function, and search point ξ i The problem setting information indicating the above is set in the problem setting calculation unit 222. In this example, the system model information includes detailed system model information and abstract system model information. The constraints include constraints on the task and constraints on the operation of the robot.

[0089] The optimal control problem (OCP) is a problem in which a target state χ is reached within a predetermined target time T as shown in equation (6). d The problem is to find a control parameter α that minimizes the evaluation value of a level set function g(x(T), ξ) that indicates the reachability of x(T). In this application, optimization includes the meaning of searching for the most appropriate value possible, and is not limited to determining an absolutely optimal value. In other words, optimization in this application does not exclude cases where the level set function or other evaluation values ​​temporarily change to more inappropriate values ​​during the processing process, or where an absolutely optimal value cannot be obtained as a processing result. Minimization also includes the meaning of searching for the smallest possible value, and is not limited to determining an absolute minimum value. A model of system 50, i.e., a system model, is given by equation (5).

[0090] The control input u to the robot 5 is the system state x(t) at time t (t is a real number between 0 and T), the control parameter α, and the target state / known task parameter value β g The low level controller π is determined by L (x(t), α, β g The control parameter α is the control output from the high-level controller π as shown in equation (1). H As mentioned above, the initial state x at time t = 0 is output from (ξ). s and the target state / known task parameter value β g The set of ξ iis equivalent to

[0091]

[0092] The problem setting calculation unit 222 sets a solution search problem that indicates task execution by the robot 5, based on the problem setting information set by the system model setting unit 221. The solution search problem indicates a problem of finding a solution that satisfies the constraint conditions set in the problem setting information. More specifically, the problem setting calculation unit 222 sets the above-mentioned optimal control problem under the constraint conditions set in the problem setting information.

[0093] The problem setting calculation unit 222 solves the set optimal control problem and calculates the smallest possible evaluation value of the level set function g as the optimal value g i * and the optimum value g i * The control parameter α i * The problem setting calculation unit 222 can calculate the search point ξ i The optimal value g calculated under i * and the control parameter α i * The problem setting calculation unit 222 outputs the optimal control solution information indicating the control parameter α i * The system state x(t) may be determined based on a system model set based on the above, and the determined system state x(t) may be included in the optimal control solution information.

[0094] The data updating unit 223 updates the search point ξ indicated in the optimal control solution information input from the problem setting calculation unit 222. i , the optimal value g i * and the control parameter α i * The high-level controller π H and the training data D of the level set function opt Update the training data D opt is the search point ξ i The optimal value for each g i * and the control parameter α i *The set of data constitutes the accumulated data set, also referred to as the acquired data set.

[0095] The learning setting unit 230 sets the acquired data set D opt The level set function g^ is learned using the above, and the prediction accuracy evaluation function J is calculated based on the level set function g^ obtained by learning. g^ and the controller learning evaluation function J h The learning setting unit 230 sets the prediction accuracy evaluation function J g^ As described above, the level set function g^ is a function for calculating an evaluation value relating to the possibility of reaching the goal state for a combination of the system state and the goal state / known task parameter values. The prediction accuracy evaluation function J g^ is the prediction accuracy of the level set function ĝ and thus the high-level controller π H This is a function for evaluating the prediction accuracy of the level set function g. The prediction accuracy is used to determine whether or not it is necessary to continue learning the level set function g. h is the high-level controller π H This is a function used as a loss function in learning.

[0096] The level set function learning unit 231 learns the acquired data set D opt The level set function g^ is learned using the above as training data. A model of the level set function g^, i.e., a function form, is set in advance in the level set function learning unit 231. In learning the level set function g^, the level set function learning unit 231 searches for parameters of the level set function g^ that minimize the evaluation function for level set function learning, for example. The evaluation function for level set function learning is, for example, as shown in equation (7), obtained data set D opt The search point ξ, which is an element of i The evaluation value g^(ξ i , α i ) and the target value g i In equation (7), the sum of the squares of the absolute values ​​of the differences between E (ξi,gi,αi)∈Dopt [...] is the element of the search point ξ i , target value g i , control parameter αi An acquired data set D including a set of opt This indicates the expected value of the evaluation value for...

[0097]

[0098] Target value g i is the search point ξ i Whether or not the value is negative depends on whether or not it is possible to reach the target state of the system state x(t) derived by the above. The level set function learning unit 231 sets the level set function g determined by learning in the prediction accuracy evaluation function setting unit 232 and the controller learning evaluation function setting unit 234.

[0099] Therefore, the level set function learning unit 231 uses the acquired data set D opt The search point ξ, which is an element of i For each state, determine whether the system state x(t) can reach the target state, and obtain the evaluation result of whether it can be reached. opt The level set function learning unit 231 may add, for example, the search point ξ at which the final system state x(T) at time T approximates the target state. i The smaller the target value g i In addition, the level set function learning unit 231 may add a search point ξ at which the system state x(t) can reach the target state. i On the other hand, the search point ξ that reaches the goal state early i The smaller the target value g i may be added.

[0100] The prediction accuracy evaluation function setting unit 232 sets the acquired data set D opt , and the prediction accuracy evaluation function J is calculated based on the level set function g set by the level set function learning unit 231. g^ The prediction accuracy evaluation function setting unit 232 sets, for example, the search point ξ i The variance σ of the evaluation value of the level set function g^ for each g^ (ξ i , α i ) is the prediction accuracy evaluation function J g^ The variance σ g^ (ξ i , α i ) is the larger the value, the more likely it is that the search point ξi Since the determination result of whether the target state can be reached or not may differ for each level set function g, it functions as an indicator that the prediction accuracy of the target state reachability by the level set function g is low. In addition, since the system state is optimized based on the level set function g, the variance σ g^ (ξ i , α i ) is the higher the value, the higher the H The prediction accuracy evaluation function setting unit 232 can also be used as an index of the prediction accuracy of the system state predicted by the set prediction accuracy evaluation function J g^ The evaluation value is set in the prediction accuracy evaluation unit 233.

[0101] The prediction accuracy evaluation unit 233 uses the prediction accuracy evaluation function J set by the prediction accuracy evaluation function setting unit 232. g^ The prediction accuracy evaluation unit 233 uses the evaluation value of the prediction accuracy evaluation function J to determine whether or not it is necessary to continue learning the level set function g. g^ The prediction accuracy evaluation unit 233 may determine whether or not learning should continue by further considering the learning conditions of the level set function g^. For example, the prediction accuracy evaluation unit 233 may determine whether or not learning should continue by further considering the learning conditions of the previous search point ξ i-1 Whether or not learning needs to continue is determined based on whether or not the amount of change in the level set function g^ at the current time point is equal to or greater than a predetermined reference amount of change from the level set function g^ obtained based on the above.

[0102] When it is determined that learning needs to continue, the prediction accuracy evaluation unit 233 outputs a learning continuation flag indicating that learning needs to continue to the search point set setting unit 210. When the learning continuation flag is input from the prediction accuracy evaluation unit 233, the search point selection unit 212 selects the acquired data set D opt A new search point ξ i Select the selected search point ξ i is output to the data acquisition unit 220. Therefore, the new search point ξ i The level set function g continues to be learned based on

[0103] If it is determined that continuation of learning is unnecessary, the prediction accuracy evaluation unit 233 does not output a learning continuation flag to the search point set setting unit 210, but causes the level set function learning unit 231 to end learning of the level set function g^, and causes the high-level controller learning unit 240 to H After that, the level set function learning unit 231 sets the level set function g^ at the time when the learning is completed in the motion planning unit 34 of the robot controller 3. Furthermore, the high-level controller learning unit 240 sets the level set function g^ at the time when the learning is completed in the motion planning unit 34 of the robot controller 3. H The parameters are set in the motion planning unit 34 and the high-level control unit 35.

[0104] The controller learning evaluation function setting unit 234 calculates the evaluation value of the level set function g^ set by the level set function learning unit 231 and the evaluation value of the high-level controller π set by the high-level controller learning unit 240. H Based on the controller learning evaluation function J h The controller learning evaluation function setting unit 234 determines the determined controller learning evaluation function J h is set in the high-level controller learning unit 240. The controller learning evaluation function setting unit 234 sets the average value μ of the level set function g^ as shown in, for example, equation (8). g^ (predicted mean value), variance μ of the level set function g g^ (prediction variance) and high-level controller π H Control output and control parameter α i Sum of squared differences with |π H (ξ i ) -α i | 2 The weighted sum of the controller learning evaluation function J h In equation (8), k and λ are the variance μ g^ , sum of squared differences |π H (ξ i ) -α i | 2 The weight parameters k and λ are each a predetermined positive real value of 0 or greater than 0.

[0105]

[0106] As shown in equation (8), the controller learning evaluation function J h The average value μ of the level set function g^ is g^ By including the high-level controller π H is learned. The controller learning evaluation function J h The variance σ of the level set function g^ is g^ By including the high-level controller π H is learned. The controller learning evaluation function J h The sum of squared differences of the components |π H (ξ i ) -α i | 2 By including the high-level controller π, it becomes easier to obtain the control parameter α as the control output. H is learned.

[0107] The high-level controller learning unit 240 learns the acquired data set D opt is referred to as training data, and the controller learning evaluation function J h Using the high-level controller π H The high-level controller learning unit 240 learns the high-level controller π H The high-level controller learning unit 240 sets a model of the high-level controller π H In the learning of the controller, for example, the controller learning evaluation function J h A high-level controller π that minimizes H The high-level controller learning unit 240 searches for the parameters of the high-level controller π H are set in the problem setting calculation unit 222 and the controller learning evaluation function setting unit 234.

[0108] Next, an example of data flow in the skill learning unit 15 according to this embodiment will be described. Fig. 11 is a data flow diagram illustrating data flow in the skill learning unit 15 according to this embodiment. The search point set initialization unit 211 of the search point set setting unit 210 uses the target parameter information stored in the storage device 2 to set the search point set Ξ checkThe search point set initialization unit 211 sets one initial state x by referring to the target parameter information, for example. si and one target state / known task parameter value β g The possible combinations of these are defined as search points ξ, and a search point set Ξ containing N different search points ξ is set. check Configure.

[0109] The search point selection unit 212 selects the search point set Ξ configured in the search point set setting unit 210. check One unprocessed search point ξ i The search point selection unit 212 randomly selects the prediction accuracy evaluation function J from the prediction accuracy evaluation function setting unit 232. g^ If is set, the unprocessed search point ξ i Among these, the prediction accuracy evaluation function J g^ The search point ξ that maximizes the evaluation value of i The search point selection unit 212 selects the selected search point ξ i to the data acquisition unit 220.

[0110] The system model setting unit 221 of the data acquisition unit 220 calculates the system model information stored in the storage device 2, other setting information, and the search point ξ input from the search point set setting unit 210. i The problem setting information is constructed using the following: system model information, constraints, low-level controller, target time, solution search function, and search point ξ i The system model setting unit 221 sets the constructed problem setting information in the problem setting calculation unit 222.

[0111] The problem setting calculation unit 222 calculates the search point ξ under the system model, constraints, and low-level controllers indicated in the problem setting information set by the system model setting unit 221. i The problem setting calculation unit 222 uses the solution search function to solve the optimal control problem as a solution search problem for i * and the optimum value g i * The control parameter α i * Calculate the search point ξi The optimal value g i * and the control parameter α i * The optimum control solution information including the set of the above is output to the data update unit 223.

[0112] The data updating unit 223 updates the search point ξ indicated in the optimal control solution information input from the problem setting calculation unit 222. i , the optimal value g i * and the control parameter α i * The acquired data set D opt Update.

[0113] The level set function learning unit 231 of the learning setting unit 230 uses the acquired data set D updated by the data updating unit 223. opt as training data, and learns the level set function g^ indicated in the preset level set function information. The level set function learning unit 231 sets the level set function g^ obtained by learning in the prediction accuracy evaluation function setting unit 232 and the controller learning evaluation function setting unit 234.

[0114] The prediction accuracy evaluation function setting unit 232 sets the acquired data set D opt , and the prediction accuracy evaluation function J is calculated based on the level set function g set by the level set function learning unit 231. g^ The prediction accuracy evaluation function setting unit 232 determines the determined prediction accuracy evaluation function J g^ The evaluation value is set in the prediction accuracy evaluation unit 233.

[0115] The prediction accuracy evaluation unit 233 determines whether or not it is necessary to continue learning the level set function g^ during learning, using the evaluation value of the prediction accuracy evaluation function Jg^ set by the prediction accuracy evaluation function setting unit 232. When determining that it is necessary to continue learning, the prediction accuracy evaluation unit 233 outputs a learning continuation flag indicating that it is necessary to continue learning to the search point set setting unit 210.

[0116] The controller learning evaluation function setting unit 234 sets the controller learning evaluation function J based on the evaluation value of the level set function g^ set by the level set function learning unit 231 and the high-level controller set by the high-level controller learning unit 240. h The controller learning evaluation function setting unit 234 determines the determined controller learning evaluation function J h is set in the high-level controller learning unit 240.

[0117] The high-level controller learning unit 240 learns the acquired data set D opt is referred to as training data, and the controller learning evaluation function J h Using the high-level controller π H The high-level controller learning unit 240 learns the high-level controller π H are set in the problem setting calculation unit 222 and the controller learning evaluation function setting unit 234.

[0118] (7) Processing Flow Next, an example of the learning process by the skill learning unit 15 according to this embodiment will be described. FIG. 12 is a flowchart illustrating the learning process according to this embodiment. (Step S101) The search point set initialization unit 211 of the search point set setting unit 210 sets the search point set Ξ using the target parameter information stored in the storage device 2. check Then, the process proceeds to loop L11, where data acquisition / learning processing is started.

[0119] (Step S102) The search point selection unit 212 selects a search point set Ξ check The search point ξ to obtain the next data from i (Step S103) The system model setting unit 221 selects the selected search point ξ i (Step S104) The problem setting calculation unit 222 solves the optimal control problem based on the problem setting information and the high-level controller, and calculates the optimal value g of the evaluation value using the level set function as the evaluation function. i * and the optimal value g i * The control parameter α i *(Step S105) The data update unit 223 calculates the search point ξ to be processed. i , the optimal value g i * , control parameter α i * The data set D is obtained by adding opt Update.

[0120] (Step S106) The level set function learning unit 231 learns the acquired data set D opt (Step S107) The controller learning evaluation function setting unit 234 learns the level set function g^ using the level set function g^ and the high-level controller π H Based on the controller learning evaluation function J h and set it in the high-level controller learning unit 240. (Step S108) The high-level controller learning unit 240 opt The controller learning evaluation function J h Using the high-level controller π H The high-level controller learning unit 240 learns the learned high-level controller π H are set in the problem setting calculation unit 222 and the controller learning evaluation function setting unit 234.

[0121] (Step S109) The prediction accuracy evaluation function setting unit 232 sets the level set function g^ and the high-level controller π based on the level set function g^. H Prediction accuracy evaluation function J for evaluating the prediction accuracy of g^ The prediction accuracy evaluation function setting unit 232 determines the determined prediction accuracy evaluation function J g^ The evaluation value of the prediction accuracy evaluation function J is set in the prediction accuracy evaluation unit 233 and the search point set setting unit 210. g^ The evaluation value of the next search point ξ i+1 This can be referenced in the processing of step S102.

[0122] (Step S110) The prediction accuracy evaluation unit 233 calculates the prediction accuracy evaluation function J g^Using the evaluation value of and predetermined learning conditions, the prediction accuracy evaluation unit 233 determines whether or not learning of the level set function g^ currently being learned needs to be continued. (Step S111) When it is determined that learning needs to be continued (YES in step S111), the prediction accuracy evaluation unit 233 outputs a learning continuation flag to the search point set setting unit 210. By proceeding to the processing of step S102, the processing of loop L11 is repeated. The prediction accuracy evaluation unit 233 determines whether or not learning of the search point ξ at that time needs to be continued. i is considered to have been processed, and the number of processed search points is incremented by 1. If the number of processed search points has not reached N, the process of loop L11 is repeated for the next search point. If the number of processed search points reaches N, the process exits from loop L11 and ends the data acquisition / learning process. If it is determined that learning should not be continued (NO in step S111), the process immediately exits from loop L11 and ends the data acquisition / learning process. Thereafter, the process of FIG. 12 ends. The level set function learning unit 231 sets the level set function g^ in the motion planning unit 34 of the robot controller 3. The high-level controller learning unit 240 learns the high-level controller π H The parameters are set in the motion planning unit 34 and the high-level control unit 35.

[0123] (8) Motion Planning The motion planning unit 34 according to this embodiment calculates the search point ξ and the control parameter π calculated based on the search point ξ in the motion planning of the robot 5. H The evaluation value g^(ξ,π) of the trained level set function g^ set for (ξ) H (ξ)) is the evaluation value g^ * The control parameter π H (ξ) is the trained high-level controller π for the search point ξ H Generally, the search point ξ related to the newly set motion plan is obtained as an output from the search point ξ used in learning the level set function g^ etc. i Since the target state may differ from the target state, the possibility of reaching the target state is unknown. The action planning unit 34 determines the possibility of reaching the target state based on the calculated evaluation value. As described above, the search point ξ is determined by the initial state x of the task to be executed. s and the target state / known task parameter value β gThe parameter set includes the evaluation value g^ as an element. * It is possible to determine whether the target state can be reached by determining whether or not is equal to or less than 0. The motion planning unit 34 further controls the high-level controller π based on the search point ξ. H The control parameter α (= π H (ξ)) satisfies a predetermined constraint, and the evaluation value g^ * is less than or equal to 0 and the constraint condition is satisfied, it may be determined that the goal state can be reached.

[0124] When the motion planning unit 34 determines that the target state can be reached, the motion planning unit 34 calculates the control parameter π obtained based on the search point ξ. H (ξ) is output as a control parameter α to the low-level control unit 36. Therefore, the movement of the robot 5 is controlled based on the search point ξ that is determined to be capable of reaching the goal state.

[0125] When determining that the target state has not been reached, the motion planning unit 34 H The sum π obtained by adding the adjustment amount Δα to (ξ) H (ξ)+Δα is defined as the adjusted control parameter α', and the evaluation value g^(ξ,α') of the level set function g^ for the search point ξ and the adjusted control parameter α' is expressed as the evaluation value g^ * The motion planning unit 34 calculates the evaluation value g^ as exemplified by equation (9). * A control parameter α′ for which is equal to or less than 0 is searched for as a control parameter that enables the target state to be reached.

[0126]

[0127] The motion planning unit 34 further uses the control parameter α′ to determine whether or not a predetermined constraint condition is satisfied, and determines whether or not to adopt the control parameter α′ obtained by the search. If it is determined not to adopt the control parameter α′, the motion planning unit 34 discards the obtained control parameter α′ and calculates the evaluation value g^ *may be set to 0 or less and a new control parameter α' that satisfies the constraint may be newly searched for. The motion planning unit 34 outputs the obtained control parameter α' to the low-level control unit 36. Therefore, even if it is determined that the search point ξ cannot reach the target state, the system state can be made to reach the target state by adjusting the control parameter that is the optimal solution for the search point ξ.

[0128] Although the search point ξ does not give a target state, and the control parameter α' is not an optimal solution for the search point ξ, the control parameter α' that makes it possible to reach the target state is searched for. For example, assume that, in the operating environment of the robot 5, another moving object enters the planned path of travel estimated by the control parameters that are the optimal solution at a certain search point ξ. Under this assumption, the motion planning unit 34 updates the search point ξ to a position sufficiently away from the predicted path of the other moving object as an element of the target state.

[0129] Thereafter, the motion planning unit 34 executes the above-described processing for the set search point ξ. That is, the motion planning unit 34 determines whether the target state can be reached for the set search point ξ and the control parameter α that provides the optimal solution. When the motion planning unit 34 determines that the target state can be reached, it derives control parameters that provide the optimal solution for the search point ξ. When the motion planning unit 34 determines that the target state cannot be reached, it adjusts the control parameters that provide the optimal solution and searches for control parameters that enable the target state to be reached. Then, after the moving object passes through the planned travel path based on the original search point ξ, the motion planning unit 34 may reset the search point ξ to include the original target state and execute the above-described processing for the reset search point ξ. Therefore, by changing the search point ξ, contact or collision with the moving object can be avoided, improving the possibility of reaching the target state.

[0130] Next, an example of a motion plan according to this embodiment will be described. FIG. 13 is a flowchart illustrating a motion plan according to this embodiment. (Step S201) The motion planning unit 34 sets a search point ξ when planning the motion of the robot 5. As described above, the search point ξ indicates the initial state, the target state, and the task parameters of the control system. (Step S202) The motion planning unit 34 sets a learned high-level controller π based on the search point ξ. H , and the high-level controller π H Control output π from H The motion planning unit 34 calculates the parameter α by using the search point ξ and the control output π. H The evaluation value g of the level set function g^(ξ,π(ξ)) for (ξ) * Calculate.

[0131] (Step S203) The operation planning unit 34 calculates the evaluation value g * is equal to or less than 0 and the motion plan for the search point ξ satisfies a predetermined constraint condition. If it is determined that it satisfies the constraint condition (YES in step S203), the process proceeds to step S204. If it is determined that it does not satisfy the constraint condition (NO in step S203), the process proceeds to step S204.

[0132] (Step S204) The motion planning unit 34 calculates the control parameter π H (ξ) as the control parameter α to the low-level control unit 36. Therefore, the motion planning unit 34 can execute control of the robot 5 using the control parameter α output from the high-level controller as it is. (Step S205) The motion planning unit 34 calculates the level set function g^(ξ,π H (ξ) + Δα) is less than or equal to 0, and the adjusted control parameter π H (ξ)+Δα is used to search for an adjustment value Δα that satisfies the constraint conditions. (Step S206) The operation planning unit 34 searches for the adjusted control parameter π obtained by the search. H The adjusted control parameter α' is output to the low-level control unit 36 ​​as (ξ)+Δα. This causes the motion planning unit 34 to control the robot 5 using the adjusted control parameter α'. Thereafter, the processing of FIG. 13 ends.

[0133] Second Embodiment Next, a second embodiment of the present invention will be described, focusing on the differences from the first embodiment. Common reference numerals are used for parts common to the first embodiment, and the same description will be used unless otherwise specified. FIG. 14 is a data flow diagram illustrating the data flow in the skill learning unit 15 according to this embodiment. The data acquisition unit 220 according to this embodiment includes a first data update unit 223-1 and a second data update unit 223-2 instead of the data update unit 223.

[0134] The problem setting calculation unit 222 uses the level set function as the evaluation function as described above to solve the optimal control problem, and calculates the minimum evaluation value as the optimal value g i * and the optimum value g i * The control parameter α i * However, the problem setting calculation unit 222 according to this embodiment calculates the optimal value g i * One or more evaluation values ​​g calculated in the calculation process until i (In this application, this may be referred to as a non-optimal solution) and their respective evaluation values ​​g i The control parameter α i The problem setting calculation unit 222 stores the set of the search point ξ i , the optimal value g i * and the control parameter α i * The problem setting calculation unit 222 outputs optimal control solution information including the set of the search point ξ to the first data update unit 223-1 and the second data update unit 223-2. i , the non-optimal value g obtained in the calculation process i and the control parameter α i The second data updating unit 223-2 outputs non-optimal control solution information including the set of the optimal control solution information and the non-optimal control solution information to the first data updating unit 223-1. Therefore, both the optimal control solution information and the non-optimal control solution information are provided to the first data updating unit 223-1. In contrast, the second data updating unit 223-2 is provided with the optimal control solution information but not with the non-optimal control solution information.

[0135] The first data updating unit 223-1 updates the search point ξ indicated in the optimal control solution information input from the problem setting calculation unit 222. i , the optimal value g i * and the control parameter α i * and the search point ξ shown in the non-optimal control solution information i , non-optimal value g i and the control parameter α i and the first data set D g The first data set D g is further search point ξ i , non-optimal value g i and the control parameter α i The acquired data set D opt The first data set D g is the acquired data set D opt Instead, the level set function ĝ is learned in the level set function learning unit 231 and the prediction accuracy evaluation function J is learned in the prediction accuracy evaluation function setting unit 232. g^ It is used to set more search points ξ i and the control parameter α i Since the level set function g^ can be learned in association with the set of , it is possible to obtain a level set function g that can more accurately explain the dependency on the control parameter α, which is the explanatory variable.

[0136] The second data updating unit 223-2 updates the search point ξ indicated in the optimal control solution information input from the problem setting calculation unit 222. i , the optimal value g i * and the control parameter α i * and accumulate the set of data D h The second data set D h is the acquired data set D opt The second data set D h is the acquired data set D opt Instead, the high-level controller π H It is used to learn.

[0137] Next, an example of the learning process by the skill learning unit 15 according to this embodiment will be described. FIG. 15 is a flowchart illustrating the learning process according to this embodiment. The process in FIG. 15 includes steps S101 to S103, S114 to S117, S107, S118, and S109 to S111. The description of FIG. 12 applies to the processes of steps S101 to S103, S107, and S109 to S111. In the process in FIG. 15, after step S103, the process proceeds to step S114.

[0138] (Step S114) The problem setting calculation unit 222 solves the optimal control problem based on the problem setting information and the high-level controller, and calculates the optimal value g of the evaluation value using the level set function as the evaluation function. i * and the optimal value g i * The control parameter α i * The problem setting calculation unit 222 calculates the optimal value g i * The non-optimal solution g calculated during the calculation process (during the calculation) until i and the non-optimal solution g i The control parameter α i Save the pair with .

[0139] (Step S115) The first data update unit 223-1 updates the search point ξ i , the optimal value g i * and the control parameter α i * and the search point ξ i , non-optimal value g i and the control parameter α i and the first data set D g (Step S116) The second data update unit 223-2 updates the search point ξ i , the optimal value g i * and the control parameter α i * The second data set D h Update.

[0140] (Step S117) The level set function learning unit 231 learns the first data set D g Then, the process proceeds to step S107. After the process of step S107, the process proceeds to step S118. (Step S118) The high-level controller learning unit 240 learns the level set function g^ using the second data set D h The controller learning evaluation function J h Using the high-level controller π H The high-level controller learning unit 240 learns the learned high-level controller π H is set in the problem setting calculation unit 222 and the controller learning evaluation function setting unit 234. Then, the process proceeds to step S109.

[0141] The first data set D to the first data update unit 223-1 g and the second data set D to the second data update unit 223-2. h and updating one or both of the level set function g or the high-level controller π H In the learning of the first data set D, experience replay (ER) may be applied. Experience replay is a technique in which multiple sets of transition information are stored in a storage area called a replay buffer (RB), one set is randomly selected from the stored transition information, and the selected sets of transition information are used sequentially for learning. One set of transition information includes information such as the states before and after a transition in one state transition, the conditions for the transition, and the amount of change in the evaluation value of the evaluation function due to the transition. g In generating the optimal solution, a data augmentation technique such as Hindsight Experience Replay may be applied using a non-optimal solution. Hindsight Experience Replay is described in, for example, the following literature: M. Andrychowicz, et al.: “Hindsight Experience Replay”, Proc of NIPS (2017)

[0142] In the above example, the second data set D h In this case, the search point ξ i , non-optimal value gi and the control parameter α i The second data set D does not include the set D. h In the search point ξ i , non-optimal value g i and the control parameter α i However, the second data set D h The search point ξ included in i , non-optimal value g i and the control parameter α i The set is the first data set D g The search point ξ included in i , non-optimal value g i and the control parameter α i It is part of the set.

[0143] <Third Embodiment> Next, the third embodiment of the present invention will be described, focusing on the differences from the first embodiment. Common reference numerals are used for parts common to the first and second embodiments, and the same explanations will be used unless otherwise specified. Figure 16 is a data flow diagram illustrating the data flow in the skill learning unit 15 according to this embodiment. The skill learning unit 15 according to this embodiment includes a system model learning unit 250.

[0144] The system model learning unit 250 acquires data indicating control results (referred to herein as "control result data") when controlling the movement of the robot 5 in a target system using a high-level controller that has been trained or is currently being trained. During control, the high-level controller uses search points that include, as elements, target parameter information related to the skill to be controlled. The control result data includes at least the search points at each time, the control parameters output from the high-level controller, and the system state that results from the control, and these are associated with each other. The control result data is associated with the target parameter information related to the skill to be controlled. During system model training, the system model learning unit 250 can determine system model parameters for the target system, using the system state at each time as output and the control parameters and target parameters indicated by the search points as input, based on the control result data. The system model learning unit 250 sets system model information indicating the parameters of the trained system model in the system model setting unit 221. The system model learning unit 250 outputs the acquired control result data to the data update unit 223.

[0145] The system model learning unit 250 may acquire an evaluation value of the level set function at each time from the system 50 as part of the control result, and include the acquired evaluation value in the control result data in association with the control parameters, search points, and system state at that time. The evaluation value of the level set function is calculated in the process of solving an optimal control problem in the motion control of the robot 5. The data updating unit 223 receives the search points and control parameters included in the control result data as input, and generates control solution information including the evaluation value as output in the acquired data set D opt In this case, the control solution information derived from the control result data is used for learning the level set function in the level set function learning unit 231 and the high-level controller in the high-level controller learning unit 240.

[0146] Next, an example of the functional configuration of the system model learning unit 250 according to this embodiment will be described. Fig. 17 is a schematic block diagram showing an example of the functional configuration of the system model learning unit 250 according to this embodiment. The system model learning unit 250 includes a high-level controller evaluation unit 250a, a skill learning low-level controller 250b, a data processing unit 250c, a data collection task management unit 250e, a transition data storage unit 250f, and a system model learning processing unit 250g.

[0147] The high-level controller evaluation unit 250a is set with the parameters of the high-level controller learned by the high-level controller learning unit 240. The high-level controller evaluation unit 250a uses the learned high-level controller to calculate control parameters based on search points derived from target parameter information set by the data collection task management unit 250e. The high-level controller evaluation unit 250a outputs the calculated control parameters to the skill learning low-level controller 250b.

[0148] The skill learning low-level controller 250b calculates a control output based on the system state collected from the system 50, the control parameters input from the high-level controller evaluation unit 250a, and the target state / known task parameter values ​​indicated in the target parameter information, and outputs the calculated control output to the system 50 to be controlled. The data processing unit 250c constructs control result data by associating the control parameters used to control the system 50, the control points acquired from the data collection task management unit 250e, the evaluation values ​​of the level set functions, and the system state acquired from the system 50. The data processing unit 250c stores the constructed control result data in a control result storage unit 250d and a transition data storage unit 250f.

[0149] The control result storage unit 250d aggregates control result data for each task. The control result data aggregated and stored in the control result storage unit 250d is read out by the data update unit 223. The control result storage unit 250d may output the acquired control result data to the data update unit 223 every time it acquires new control result data. The data collection task management unit 250e accumulates control result data for each time, and the accumulated control result data is accumulated and formed as transition data for each skill.

[0150] The system model learning processing unit 250g uses the transition data stored in the transition data storage unit 250f to learn a system model indicated by preset system model information. The system model learning processing unit 250g receives the control parameters indicated in the transition data and the target parameter information indicated at the search points as input, and determines, for each skill, system model parameters for estimating the system state at each time of the target system as output. The system model learning processing unit 250g sets, in the system model setting unit 221, system model information indicating the learned system model.

[0151] The system model learning processor 250g can use a system model exemplified by Equation (5). By setting a small unit time in this system model, the time derivative of the system state exemplified on the left side of Equation (5) can approximate the amount of change from the system state at the current time to the system state at the time unit time after the current time (i.e., the next time). Therefore, the system model learning processor 250g can use, as the system model, a function that outputs the system state at the next time and inputs the system state at the current time, control parameters, and target parameter information.

[0152] Furthermore, the system model learning unit 250 may learn the system model separately from the learning of the level set function and the high-level controller, or may learn the system model for each search point related to the learning. As described above, the search point includes target parameter information as an element. In learning the system model, a preset search point set Ξ checkInstead of search points belonging to the above group, search points that provide target parameter information for use in operating the system 50 may be used.

[0153] Next, an example of the learning process by the learning device 1 according to this embodiment will be described. FIG. 18 is a flowchart illustrating the learning process according to this embodiment. The processes of steps S301 and S302 are executed independently of the learning of the level set function and the high-level controller. The processes of steps S304 and S305 are executed independently of the learning of the search point ξ i This is performed in association with the level set function and high-level controller training according to step S303.

[0154] (Step S301) The system model learning unit 250 randomly acquires control result data from the system 50 operating in the real environment. Random acquisition includes acquiring control result data at randomly determined periods, randomly determining whether acquisition is necessary for each operation, and acquiring data only when it is determined that acquisition is necessary. One or both of the skills and the target parameters may differ for each operation. (Step S302) The system model learning unit 250 learns a system model for each skill based on the acquired control result data. The system model learning unit 250 sets system model information indicating the parameters of the learned system model in the system model setting unit 221. The system model learning unit 250 outputs the control result data acquired during the system model learning process to the data update unit 223. Then, the process proceeds to loop L31, where the data acquisition / learning process begins. Note that, although the number of search points is N in the example of FIG. 18 , the process in loop L31 may be executed for each operation of the system 50. At least one item of the initial state, task parameters, and target parameter information that form the search points may be different for each run.

[0155] (Step S303) The skill learning unit 15 learns the level set function and the high-level controller. The processing of this step may be the same as the processing of steps S102 to S111 of the loop L11 illustrated in Fig. 12. However, when it is determined in step S111 that learning should not be continued (step S111 NO), the prediction accuracy evaluation unit 233 proceeds to the processing of step S303 without changing the number of search points that have been processed.

[0156] (Step S304) The system model learning unit 250 acquires control result data from the system 50 that controlled the operation of the robot 5 in the real environment using the trained high-level controller. (Step S305) The system model learning unit 250 learns a system model based on the acquired control result data. The system model learning unit 250 sets system model information indicating the parameters of the trained system model in the system model setting unit 221. The system model learning unit 250 outputs the control result data acquired in the system model learning process to the data updating unit 223. The prediction accuracy evaluation unit 233 calculates the search point ξ at that time. i is considered to have been processed, and the number of processed search points is incremented by 1. If the number of processed search points has not reached N, the process of loop L31 is repeated for the next search point. When the number of processed search points reaches N, the process exits from loop L31 and the data acquisition / learning process ends. Then, the process of FIG. 18 ends.

[0157] The above description has mainly focused on the differences from the first embodiment, but instead of the data update unit 223, similar to the skill learning unit 15 according to the second embodiment, the skill learning unit 15 according to this embodiment may include a first data update unit 223-1 and a second data update unit 223-2. In this case, the first data set D generated by the first data update unit 223-1 g is the learning of the level set function g^ in the level set function learning unit 231 and the prediction accuracy evaluation function J in the prediction accuracy evaluation function setting unit 232. g^ The second data set D generated by the second data update unit 223-2 is used to set h is the high-level controller π HThe process of step S303 illustrated in Fig. 18 is similar to the processes of steps S102, S103, S114 to S117, and S107 to S111 of loop L21 illustrated in Fig. 15.

[0158] The control result data generated by the system model learning unit 250 is used for learning the level set function ĝ and the prediction accuracy evaluation function J g^ In this case, the control result data generated by the system model learning unit 250 can be used to set the high-level controller π H It does not have to be used for learning.

[0159] <Minimum Configuration Example> Next, an example of the minimum configuration of an embodiment of the present application will be described. Fig. 19 is a schematic block diagram showing an example of the minimum configuration of a learning device 1 according to an embodiment of the present application. The learning device 1 includes: a level set function learning unit 231 that receives as input motion parameters that define the motion target and motion environment of the robot and control parameters of the robot, and learns a level set function that outputs an evaluation value related to the attainment of the motion target based on the motion parameters and the control parameters; and a high-level controller learning unit 240 that learns, based on the prediction accuracy of the level set function and the control parameters, a high-level controller that determines control parameters for realizing a target motion based on the motion parameters of the robot.

[0160] With this configuration, the high-level controller learning unit 240 learns a high-level controller for determining control parameters for realizing a target motion based on motion parameters used to control the robot, and the level set function learning unit 231 learns a level set function for determining an evaluation value relating to the achievability of a motion goal based on the motion parameters and the control parameters. Therefore, the achievability of a motion goal based on the motion parameters can be determined using an evaluation value obtained from the motion parameters related to the control of the robot using the learned level set function. This makes it possible to efficiently plan the motion of the robot using the motion parameters.

[0161] 20 is a schematic block diagram showing an example of the minimum configuration of a control device 3′ according to an embodiment of the present application. The control device 3′ includes a high-level control unit 35 that determines control parameters for realizing a target motion based on motion parameters that define the motion target and motion environment of the robot, and a motion planning unit 34 that calculates an evaluation value indicating the feasibility of the target motion based on the motion parameters and control parameters using a level set function, and uses the control parameters to control the motion of the robot when it is determined that the target motion is feasible based on the evaluation value, and searches for control parameters that make the target motion feasible based on the evaluation value when it is determined that the target motion is not feasible based on the evaluation value.

[0162] With this configuration, the high-level control unit 35 obtains control parameters based on the motion parameters, and the motion planning unit 34 evaluates the feasibility of a target motion using the motion parameters based on the motion parameters and evaluation values ​​obtained from the control parameters. If it is determined that the target motion is feasible, the control parameters obtained by the high-level control unit 35 are used for motion control, and if it is determined that the target motion is not feasible, a search is made for control parameters that will make the target motion feasible based on the evaluation values. This makes it possible to promote motion planning so that the target motion using the motion parameters can be realized as much as possible.

[0163] As described above, an embodiment of the present application may be realized as a control system 100 including a learning device 1, a robot 5, and a high-level controller. An embodiment of the present application may be realized as a program for causing a computer to function as the learning device 1, or as a non-transitory storage medium that stores the program and is readable by the computer. The computer may be configured to include a processor or other integrated circuit and the non-transitory storage medium, and may be capable of executing processes instructed by instructions that constitute the program, thereby realizing the functions of the learning device 1.

[0164] The learning device 1 also includes a search point selection unit 212 that selects one search point from a search point set including a predetermined number of search points based on a prediction accuracy evaluation function that indicates the prediction accuracy of the level set function, where the search point is a pair of an operation parameter and a control parameter. With this configuration, the prediction accuracy of the level set function is evaluated for a known pair of an operation parameter and a control parameter, and the operation parameter and the control parameter to be used in learning the level set function are selected based on the evaluated prediction accuracy. By preferentially selecting operation parameters and control parameters that have low prediction accuracy and therefore a high convergence rate in the learning process, the learning of the level set function can be made more efficient.

[0165] The learning device 1 may also include a prediction accuracy evaluation unit 233 that uses a prediction accuracy evaluation function to determine whether or not it is necessary to continue learning the level set function and the high-level controller. With this configuration, the need for continued learning is quantitatively determined based on the prediction accuracy evaluated during the learning process.

[0166] Alternatively, the level set function learning unit 231 may learn the level set function using first learning data, and the high-level controller learning unit 240 may learn the high-level controller learning unit 240 using second learning data. The first learning data includes an optimal solution set that receives as input operational parameters and control parameters that provide an optimal solution for the evaluation value and outputs an optimal solution, and a non-optimal solution set that receives as input operational parameters and control parameters that provide a non-optimal solution different from the optimal solution and outputs a non-optimal solution, and the second learning data includes an optimal solution set. According to this configuration, the level set function is learned using first learning data that further outputs a non-optimal solution and includes a data set that receives as input operational parameters and control parameters that provide the non-optimal solution. A wider range of control parameters is referenced in learning the level set function that regresses the solution to the optimal control problem. Consequently, the high-level learning device is controlled to output stable control parameters even when the system state deviates from the optimal state.

[0167] The learning device 1 may include a system model learning unit 250 that uses control result data including a set of a system state obtained by controlling the operation of the robot using control parameters based on the operation parameters and the control parameters to learn a system model that has the operation parameters and control parameters as input and the system state as output. According to this configuration, the system model is learned using the control result data so that the system state actually obtained for the operation parameters and control parameters can be estimated. By updating the control problem based on the learned system model, it is possible to learn a high-level controller to adapt to the real system environment.

[0168] The high-level controller learning unit may learn the high-level controller using the operation parameters related to the control result data as input and the control parameters as output. With this configuration, the high-level controller is trained using the operation parameters actually used for control as input so that the control parameters actually obtained by control can be obtained. Therefore, the high-level controller is trained to adapt to the actual system environment.

[0169] Alternatively, a program for executing all or part of the processes performed by the learning device 1, robot controller 3, and robot 5 may be recorded on a computer-readable recording medium, and the program temporarily or non-temporarily recorded on the recording medium may be loaded into a computer system and executed to perform the processes of each component. The term "computer system" as used herein includes hardware such as an OS and peripheral devices. The term "computer-readable recording medium" refers to portable media such as flexible disks, optical magnetic disks, ROMs (Read Only Memory), and CD-ROMs (Compact Disc Read Only Memory), as well as storage devices such as hard disks built into the computer system. The program may be for implementing part of the functions described above, or may be capable of implementing the functions described above in combination with a program already recorded on the computer system.

[0170] Although the embodiments of the present application have been described in detail above with reference to the drawings, the specific configuration is not limited to this embodiment, and includes designs within the scope of the gist of the present invention.

[0171] The embodiments of the present application can be realized as a learning device, a control device, a control system, a learning method, and a storage medium.

[0172] 1... learning device, 2... storage device, 3... robot controller, 3'... control device, 4... measurement device, 5... robot, 11... processor, 12... memory, 13... interface, 14... abstract system model setting unit, 15... skill learning unit, 16... skill tuple generation unit, 31... processor, 32... memory, 33... interface, 50... system, 100... control system, 210... search point set setting unit, 211... search point set initialization unit, 212... search point selection unit, 220... data acquisition unit, 221... system model setting unit, 222... problem setting calculation unit, 223 (223-1, 223-2)... data update unit, 230... learning setting unit, 231... level set function learning unit, 232... prediction accuracy evaluation function setting unit, 233... prediction accuracy evaluation unit, 234... evaluation function setting unit for controller learning, 240... high-level controller learning unit, 250... system model learning unit

Claims

1. a level set function learning unit that learns a level set function that receives as input operation parameters that define an operation goal and an operation environment of a robot and control parameters of the robot, and outputs an evaluation value regarding the possibility of achieving the operation goal based on the operation parameters and the control parameters; a high-level controller learning unit that learns a high-level controller that determines control parameters for realizing a target motion based on the motion parameters of the robot based on the level set function and prediction accuracy of the control parameters. Learning device.

2. From a search point set including a plurality of predetermined search points, a search point selection unit that selects one search point based on a prediction accuracy evaluation function that indicates the prediction accuracy of the level set function; The search point is a pair of the operation parameter and the control parameter. The learning device according to claim 1 .

3. A prediction accuracy evaluation unit is provided that uses the prediction accuracy evaluation function to determine the need for continued learning of the level set function and the high-level controller. The learning device according to claim 2 .

4. the level set function learning unit learns the level set function using first learning data; the high-level controller learning unit learns the high-level controller learning unit using second learning data; the first learning data includes an optimal solution set that receives the operation parameters and the control parameters that provide an optimal solution for the evaluation value as input and includes the optimal solution as output, and a non-optimal solution set that receives the operation parameters and the control parameters that provide a non-optimal solution different from the optimal solution as input and includes the non-optimal solution as output, The second learning data includes the optimal solution set. The learning device according to claim 2 .

5. and a system model learning unit that uses control result data including a set of a system state obtained by controlling the operation of the robot using the control parameters based on the operation parameters and the control parameters to learn a system model that uses the operation parameters and the control parameters as inputs and outputs the system state. The learning device according to claim 1 .

6. 6. The learning device according to claim 5, wherein the high-level controller learning unit learns the high-level controller using the operation parameters related to the control result data as input and the control parameters as output.

7. The robot; the high level controller; The learning device according to claim 1 . Control system.

8. A computer provided in a learning device, A function for learning a level set function that receives as input operation parameters that define the operation goal and operation environment of the robot and control parameters of the robot, and outputs an evaluation value regarding the possibility of achieving the operation goal based on the operation parameters and the control parameters; and a function of learning a high-level controller that determines control parameters for realizing a target motion based on the motion parameters of the robot, based on the level set function and the prediction accuracy of the control parameters; A program to achieve this.

9. A learning method for a learning device, comprising: The learning device a level set function learning step of learning a level set function that receives as input motion parameters defining a motion goal and a motion environment of the robot and control parameters of the robot, and outputs an evaluation value regarding the possibility of achieving the motion goal based on the motion parameters and the control parameters; a high-level controller learning step of learning a high-level controller that determines control parameters for realizing a target motion based on the motion parameters of the robot, based on the level set function and prediction accuracy of the control parameters. How to learn.

10. a high-level control unit that determines control parameters for realizing a target motion based on motion parameters that define the motion target and motion environment of the robot; calculating an evaluation value indicating the feasibility of a target operation based on the operation parameters and the control parameters using a level set function; When it is determined that the target motion is achievable based on the evaluation value, the control parameter is used for motion control of the robot; and a motion planning unit that, when determining whether or not the target motion is realizable based on the evaluation value, searches for control parameters that make the target motion realizable based on the evaluation value. Control device.