Learning device, control device, learning method and program

The learning device optimizes the learning of modularized robot skills by evaluating generalization errors and deciding when to continue learning, addressing inefficiencies in existing systems by determining meta-parameter values effectively.

JP7806879B2Active Publication Date: 2026-01-27NEC CORP
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
JP2024504055
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Filing Date
2022-03-01
Publication Date
2026-01-27
Estimated Expiration
2042-03-01

AI Technical Summary

Technical Problem

Existing systems struggle with efficiently learning modularized robot behavior skills, leading to unnecessary learning and inefficiencies when accommodating differences in modules using meta-parameter values.

Method used

A learning device and method that includes a metaparameter learning mechanism to determine the probability distribution of parameters based on training data, evaluates generalization error, and decides whether to continue learning based on an evaluation value, integrating judgments across multiple models.

Benefits of technology

Enables efficient learning by determining when to stop or continue learning meta-parameter values, optimizing the learning process and avoiding unnecessary computations.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007806879000037
    Figure 0007806879000037
  • Figure 0007806879000038
    Figure 0007806879000038
  • Figure 0007806879000039
    Figure 0007806879000039
Patent Text Reader

Abstract

This learning device comprises: a meta-parameter learning means for learning values of a meta-parameter indicating, in a learning model in which the values of a parameter follow a probability distribution, the probability distribution on the basis of training data indicating input and output in the learning model; a generalization error evaluation means for calculating an evaluation value indicating evaluation of a generalization error of the learning model; and a learning continuation determination means for determining, on the basis of the evaluation value, the necessity of a learning continuation of the values of the meta-parameter.
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] The present invention relates to a learning device, a control device, a learning method, and program Regarding. [Background technology]

[0002] When controlling a robot required to execute a task, a system has been proposed in which robots are controlled by providing skills that modularize the robot's movements. For example, Patent Document 1 discloses a technology in which, in a system in which an articulated robot executes a given task, robot skills that can be selected depending on the task are defined as tuples, and the parameters included in the tuples are updated through learning. [Prior art documents] [Patent documents]

[0003] [Patent Document 1] International Publication No. 2018 / 219943 Summary of the Invention [Problem to be solved by the invention]

[0004] When learning modularized robot behavior skills, if differences in modules can be accommodated by learning the meta-parameter values ​​of the learning model, it will be possible to have the robot perform multiple skills using a single model. In this way, when learning the meta parameter values ​​of a learning model, if it is possible to determine whether or not learning needs to continue, it is expected that unnecessary learning can be avoided and learning can be carried out efficiently.

[0005] An example of the object of this disclosure is to provide a learning device, a control device, a learning method, and program The purpose is to provide [Means for solving the problem]

[0006] According to a first aspect of the present invention, a learning device includes: a metaparameter learning means for learning values ​​of metaparameters indicating a probability distribution in a learning model in which the values ​​of the parameters follow a probability distribution, based on training data indicating inputs and outputs in the learning model; a generalization error evaluation means for calculating an evaluation value indicating an evaluation of a generalization error of the learning model; and a learning continuation determination means for determining whether or not learning of the values ​​of the metaparameters needs to be continued based on the evaluation value. a learning continuation judgment integration means for judging whether or not learning of the meta parameter values ​​needs to be continued for the plurality of learning models as a whole, based on judgment results of the plurality of learning continuation judgment means corresponding to the plurality of learning models; Equipped with.

[0007] According to a second aspect of the present invention, the control device includes a control means for controlling the robot in accordance with the shape of the object to be grasped so that the robot grasps each of the objects to be grasped having different shapes.

[0008] According to a third aspect of the present invention, a learning method includes a computer learning a value of a meta parameter indicating a probability distribution in a learning model in which the value of the parameter follows a probability distribution, based on training data indicating an input and an output in the learning model, calculating an evaluation value indicating an evaluation of a generalization error of the learning model, and determining whether or not to continue learning the value of the meta parameter based on the evaluation value. death , determining whether or not to continue learning the values ​​of the meta parameters for all of the learning models based on a determination result of whether or not to continue learning each of the values ​​of the meta parameters corresponding to the learning models; This includes:

[0009] According to a fourth aspect of the present invention, a program causes a computer to: learn values ​​of meta parameters indicating a probability distribution in a learning model in which the values ​​of the parameters follow a probability distribution, based on training data indicating inputs and outputs in the learning model; calculate an evaluation value indicating an evaluation of a generalization error of the learning model; and determine whether or not it is necessary to continue learning the values ​​of the meta parameters based on the evaluation value. determining whether or not to continue learning the meta parameter values ​​for all of the plurality of learning models based on a determination result of whether or not to continue learning each of the plurality of meta parameter values ​​corresponding to the plurality of learning models; This is a program for executing the above. [Effects of the Invention]

[0010] According to the present invention, when learning meta parameter values ​​of a learning model, it is possible to determine whether or not learning needs to be continued. [Brief explanation of the drawings]

[0011] [Figure 1] 1 is a diagram illustrating an example of the configuration of a control system according to a first embodiment. [Figure 2] FIG. 4 is a diagram showing an example of known task parameters according to the first embodiment. [Figure 3] FIG. 4 is a diagram illustrating an example of unknown task parameters according to the first embodiment. [Figure 4] FIG. 2 is a diagram illustrating an example of a hardware configuration of a learning device according to the first embodiment. [Figure 5] FIG. 2 is a diagram illustrating an example of a hardware configuration of a robot controller according to the first embodiment. [Figure 6] 1 is a diagram illustrating a robot that grasps an object according to a first embodiment and an object to be grasped in real space. FIG. [Figure 7] FIG. 7 is a diagram showing the state shown in FIG. 6 in an abstract space. [Figure 8] FIG. 2 is a diagram illustrating an example of the configuration of a control system related to the execution of skills in the first embodiment. [Figure 9] FIG. 2 is a diagram illustrating an example of a functional configuration of a learning device related to updating of a skill database according to the first embodiment. [Figure 10] FIG. 2 is a diagram illustrating an example of the configuration of a skill learning unit according to the first embodiment. [Figure 11] 3 is a diagram showing an example of data input / output in a skill learning unit according to the first embodiment. FIG. [Figure 12] FIG. 10 is a diagram illustrating an example of a skill database update process performed by the learning device according to the first embodiment. [Figure 13] FIG. 10 is a diagram showing an example of data input / output in a skill learning unit according to the second embodiment. [Figure 14] FIG. 10 is a diagram illustrating an example of a skill database update process performed by a learning device according to a second embodiment. [Figure 15] FIG. 11 is a diagram illustrating an example of the configuration of a skill learning unit according to the third embodiment. [Figure 16]FIG. 11 is a diagram showing an example of data input / output in a skill learning unit according to the third embodiment. [Figure 17] FIG. 11 is a diagram illustrating an example of the configuration of a meta parameter processing unit according to the third embodiment. [Figure 18] 13A and 13B are diagrams illustrating an example of input and output of data in a meta parameter processing unit according to the third embodiment. [Figure 19] FIG. 11 is a diagram illustrating a first example of the configuration of a meta parameter individual processing unit according to the third embodiment. [Figure 20] 20 is a diagram showing an example of input and output of data in the meta parameter individual processing unit shown in FIG. 19. FIG. [Figure 21] FIG. 11 is a diagram illustrating a second example of the configuration of a meta parameter individual processing unit according to the third embodiment. [Figure 22] 22 is a diagram showing an example of input and output of data in the meta parameter individual processing unit shown in FIG. 21. FIG. [Figure 23] FIG. 11 is a diagram illustrating an example of a skill database update process performed by a learning device according to a third embodiment. [Figure 24] FIG. 11 is a diagram illustrating an example of a process in which a meta parameter processing unit according to the third embodiment calculates meta parameter values ​​of a predictor. [Figure 25] FIG. 11 is a diagram illustrating a first example of processing in which a meta parameter individual processing unit according to the third embodiment calculates a meta parameter value for each predictor and determines whether or not learning of the meta parameter value needs to be continued. [Figure 26] FIG. 11 is a diagram illustrating a second example of processing in which a meta parameter individual processing unit according to the third embodiment calculates a meta parameter value for each predictor and determines whether or not learning of the meta parameter value needs to be continued. [Figure 27] FIG. 10 is a diagram illustrating an example of the configuration of a learning device according to a fourth embodiment. [Figure 28] FIG. 13 is a diagram illustrating an example of the configuration of a control device according to a fifth embodiment. [Figure 29] FIG. 13 is a diagram illustrating an example of a processing procedure in a learning method according to a sixth embodiment. DETAILED DESCRIPTION OF THE INVENTION

[0012] Hereinafter, embodiments of the present invention will be described, but the following embodiments do not limit the invention according to the claims. Furthermore, not all of the combinations of features described in the embodiments are necessarily essential to the solution of the invention. For the sake of convenience, in this specification, a character with an arbitrary symbol "x" above the arbitrary letter "A" will be referred to as "A". x " is expressed as ".

[0013] First Embodiment (1) System configuration Fig. 1 is a diagram showing an example of the configuration of a control system according to the first embodiment. In the configuration shown in Fig. 1, the control system 100 includes a learning device 1, a storage device 2, a robot controller 3, a measuring device 4, and a robot 5. The learning device 1 communicates data with the storage device 2 via a communication network or by direct wireless or wired communication. The robot controller 3 also communicates data with the storage device 2, the measuring device 4, and the robot 5 via a communication network or by direct wireless or wired communication.

[0014] The learning device 1 learns the behavior of the robot 5 for executing a given task by machine learning such as self-supervised learning (SSL), and also learns a set of states in which the behavior to be learned can be executed.

[0015] However, the object for which the learning device 1 learns the behavior is not limited to a specific object, and can be any control object that is controllable and whose control can be learned. Furthermore, the behavior of the control object, such as the robot 5, is not limited to one that involves a change in position. For example, the robot 5 may acquire sensor measurement data using a sensor as one of its behaviors. The same applies to the following embodiments.

[0016] The state here refers to the state of the target system including the robot 5 and the operating environment of the robot 5. The robot 5 and the operating environment of the robot 5 are collectively referred to as the target system, or simply as the system. When a task involves handling an object, such as a task of grasping an object, the object of the task is also included in the target system.

[0017] The state of a target system is called the system state, or simply the state. The system state at the time of task completion, as defined in a task, is also called the goal state of that task, or simply the goal state. Reaching the goal state of a task is also called accomplishing the task or succeeding in the task. When a task is accomplished by executing a skill, the state at the end of skill execution corresponds to the goal state. The system state at the start of a task is also called the initial state of the task.

[0018] The learning device 1 learns skills that modularize specific actions of the robot 5. In the embodiment, a task is assumed to be achieved by executing one skill for each task, and the learning device 1 learns skills to achieve the task.

[0019] On the other hand, the robot controller 3 may execute a task by combining multiple skills. For example, the robot controller 3 may divide a given task into subtasks corresponding to skills, and combine skills for executing each subtask to plan the execution of the given task.

[0020] When learning a skill, the learning device 1 also learns a set of states in which the skill can be executed. The learning device 1 registers information about the learned skill in a skill database stored in the storage device 2. The information registered in the skill database is also called a skill tuple. The skill tuple includes various information required to execute the operation to be modularized. The learning device 1 generates a skill tuple based on detailed system model information, low-level controller information, and target parameter information stored in the storage device 2.

[0021] The storage device 2 stores information referenced by the learning device 1 and the robot controller 3. The storage device 2 stores, for example, detailed system model information, low-level controller information, target parameter information, and a skill database. The storage device 2 may be an external storage device such as a hard disk connected to or built into the learning device 1 or the robot controller 3, a storage medium such as a flash memory, or a server device that communicates data with the learning device 1 and the robot controller 3. The storage device 2 may also be composed of multiple storage devices, and may have the above-mentioned storage units distributed among them.

[0022] Detailed system model information is information that represents a model of a target system in real space. A model of a target system in real space is also called a detailed system model. To distinguish it from an abstract system model, which is an abstraction of a detailed system model, it is referred to as a "detailed" system model. The detailed system model information may be represented by a differential or difference equation representing the detailed system model. Alternatively, the detailed system model may be configured as a simulator that simulates the operation of the robot 5.

[0023] The low-level controller information is information about a low-level controller that generates an input to control the actual movement of the robot 5 based on parameter values ​​output by the high-level controller. For example, when the high-level controller generates a trajectory for the robot 5, the low-level controller may generate a control input that follows the movement of the robot 5 according to the trajectory. For example, the low-level controller may control the robot 5 by servo control using PID (Proportional Integral Differential) based on parameters output by the high-level controller.

[0024] The target parameter information is provided for each skill that the learning device 1 learns, and includes, for example, initial state information, target state / known task parameter information, unknown task parameter information, execution time information, and general constraint information. Here, the variable parts of a task are called task parameters.

[0025] Among the task parameters, those that are expressed as numerical values ​​are referred to as known task parameters. Examples of known task parameters include, but are not limited to, the size of an object in a task, such as the size of an object to be grasped if the task is to grasp an object, and the trajectory of the robot 5 for performing the task. Known task parameters can also be treated as parameters in skills, and are an example of skill parameters.

[0026] Fig. 2 is a diagram showing examples of known task parameters. Fig. 2 shows an example in which the robot 5 executes a task of grasping a cylindrical object. In this case, the radius and height of the cylindrical object correspond to examples of known task parameters.

[0027] On the other hand, among task parameters, those that are difficult to express numerically are referred to as unknown task parameters. Examples of unknown task parameters include, but are not limited to, the shape of an object in a task, such as the shape of an object to be grasped when the task is to grasp an object, and the type of movement of the robot 5 to execute the task, such as the skills required to execute the task.

[0028] Fig. 3 is a diagram showing examples of unknown task parameters. Fig. 3 shows an example in which the robot 5 executes a task of grasping objects of various shapes. In this case, the shapes of the objects correspond to an example of unknown parameters.

[0029] Furthermore, it is assumed that the control system 100 handles the system state as a numerical value, and the target state is expressed as a numerical value. For example, in the case of a task in which the robot 5 performs pick and place, the target state may be expressed as the coordinates of the target object being within a predetermined range.

[0030] The initial state information indicates a set of states in which a target skill can be executed. The state at the start of skill execution is also referred to as the initial state of the skill, or simply as the initial state. The set of initial states is also referred to as the initial state set. Let the initial state be x s or x si Here, "i" is a positive integer that represents an identification number that identifies the initial state. The time of the initial state is sometimes set to 0, and the initial state is sometimes represented as x0.

[0031] The goal state / known task parameter information is information that indicates a set of combinations of possible values ​​of a goal state, which is a state that can be reached by executing a target skill, and possible values ​​of known task parameters that are treated as explicit parameters of the target skill. For example, in the case of a skill in which the robot 5 grasps an object, the possible values ​​of the goal state may include information on stable grasping conditions such as form closure and force closure. The combination of the goal state and the known task parameter value is called the goal state / known task parameter value, and β g or β gi Here, "i" is a positive integer representing an identification number that identifies the target state / known task parameter value.

[0032] By treating the differences in the goal states of a task and the differences in the known task parameter values ​​as parameters in a skill, tasks with different goal states, known task parameter values, or both can be executed with a single skill.

[0033] For example, when the learning device 1 performs processing related to skill learning using a predictor, a target state and known task parameter values ​​can be input to the predictor to obtain an output value corresponding to the target state and known task parameter values. Here, the predictor is configured using a learning model (a model in machine learning), such as a neural network or a Gaussian process (GP).

[0034] Note that some skills may not have known task parameters. In this case, the goal state / known task parameter information may be configured as a set of values ​​that the goal state can take. Also, the goal state / known task parameter value β g may indicate the target state.

[0035] The unknown task parameter information is information related to unknown task parameters. For example, as will be described later in a third embodiment, the unknown task parameter information may indicate a probability distribution of data related to the unknown parameters. When one skill has multiple unknown task parameters, the unknown task parameter information may indicate information related to each unknown task parameter. In the first and second embodiments, a description will be given of how to handle target state / known task parameter information. In the first and second embodiments, values ​​corresponding to unknown task parameters may be indicated as fixed values.

[0036] The unknown task parameter value is set to τ or τ j Here, "j" is a positive integer representing an identification number that identifies the unknown task parameter value. Although it is difficult to express unknown task parameters numerically because it is difficult to quantify their values ​​systematically, it is possible to determine whether the unknown task parameter values ​​are the same. For example, if the unknown task parameter represents the shape of an object, it is possible to determine whether the unknown task parameter values ​​are the same by comparing the shapes of two objects. The control system 100 treats two tasks as the same task if the unknown task parameter values ​​in the two tasks are the same, and treats the two tasks as separate tasks if the unknown task parameter values ​​are different. j The "j" above can also be thought of as a positive integer that represents an identification number that identifies the task.

[0037] The execution time information is information about the time limit for skill execution. For example, the execution time information may indicate the skill execution time (the time it takes to execute the skill), or the allowable condition value for the time from the start to the end of skill execution, or both. The general constraint information is information indicating general constraint conditions, such as conditions relating to limits on the range of movement of the robot 5, limits on speed, and limits on input.

[0038] The skill database is a database of skill tuples prepared for each skill. A skill tuple may include information about a high-level controller for executing the target skill, information about a low-level controller for executing the target skill, and information about a set of combinations of states (initial states for the skill) and target states / known task parameter values ​​that allow the target skill to be executed. A set of states and target states / known task parameter values ​​that allow the target skill to be executed is also referred to as an executable state set.

[0039] The feasible state set may be defined in an abstract space that abstracts the actual space. The feasible state set can be expressed using, for example, a level set function estimated by Gaussian Process Regression (GPR) or Level Set Estimation (LSE), or an approximation function of the level set function. In other words, whether the feasible state set includes a combination of a certain state and a target state / known task parameter value can be determined by whether the value (e.g., the mean value) of the Gaussian process regression for the combination of the certain state and the target state / known task parameter value, or the value of the approximation function for the combination of the certain state and the target state / known task parameter value, satisfies a constraint for determining feasibility. In the following, an example will be described in which a level set function is used as a function indicating a set of feasible states, but the present invention is not limited to this.

[0040] After the learning process by the learning device 1, the robot controller 3 formulates an operation plan for the robot 5 based on the measurement signals supplied by the measurement device 4, the skill database, etc. The robot controller 3 generates control commands (control inputs) for causing the robot 5 to execute the planned operation, and supplies the control commands to the robot 5.

[0041] For example, the robot controller 3 converts a task to be executed by the robot 5 into a sequence for each time step (time interval) of the task that the robot 5 can accept. Then, the robot controller 3 controls the robot 5 based on a control command corresponding to an execution command of the generated sequence. The control command corresponds to a control input output by a low-level controller.

[0042] The measurement device 4 is, for example, one or more sensors such as a camera, a range sensor, a sonar, or a combination thereof that detects the state within the workspace where the robot 5 executes a task. The measurement device 4 supplies the generated measurement signal to the robot controller 3. The measurement device 4 may be a self-propelled or flying sensor (including a drone) that moves within the workspace. The measurement device 4 may also include sensors provided on the robot 5 and sensors provided on other objects within the workspace. The measurement device 4 may also include a sensor that detects sound within the workspace. In this way, the measurement device 4 may include various sensors that detect the state within the workspace and are provided at any location.

[0043] The robot 5 performs work related to a specified task based on a control command supplied from the robot controller 3. The robot 5 is a robot that operates, for example, in various factories such as an assembly factory or a food factory, or in a logistics site. The robot 5 may be a vertical articulated robot, a horizontal articulated robot, or any other type of robot. The robot 5 may supply a status signal indicating the status of the robot 5 to the robot controller 3. This status signal may be an output signal from a sensor that detects the status (position, angle, etc.) of the entire robot 5 or a specific part such as a joint, or may be a signal indicating the progress of the operation of the robot 5.

[0044] 1 is an example, and various modifications may be made to the configuration. For example, the robot controller 3 and the robot 5 may be integrated. In another example, at least two of the learning device 1, the storage device 2, and the robot controller 3 may be integrated. Furthermore, the control target of the control system 100 is not limited to a robot. The control system 100 can control various control targets that the learning device 1 can learn to control.

[0045] (2) Hardware configuration 4 is a diagram showing an example of the hardware configuration of the learning device 1. The learning device 1 includes, as hardware, a processor 11, a memory 12, and an interface 13. The processor 11, the memory 12, and the interface 13 are connected via a data bus 10.

[0046] The processor 11 functions as a controller (arithmetic unit) that performs overall control of the learning device 1 by executing a program stored in the memory 12. The processor 11 is, for example, a processor such as a CPU (Central Processing Unit), a GPU (Graphics Processing Unit), or a TPU (Tensor Processing Unit). The processor 11 may be composed of multiple processors. The processor 11 is an example of a computer.

[0047] Memory 12 is composed of various types of volatile and non-volatile memory, such as RAM (Random Access Memory), ROM (Read Only Memory), and flash memory. Memory 12 also stores programs for executing the processes performed by learning device 1. Some of the information stored in memory 12 may be stored in one or more external storage devices (e.g., storage device 2) that can communicate with learning device 1, or may be stored in a storage medium that is detachable from learning device 1.

[0048] Interface 13 is an interface for electrically connecting study device 1 to other devices. These interfaces may be wireless interfaces such as network adapters for wirelessly transmitting and receiving data to other devices, or hardware interfaces for connecting to other devices via cables or the like. For example, interface 13 may interface with input devices that accept user input (external input), such as touch panels, buttons, keyboards, and voice input devices, display devices such as displays and projectors, and sound output devices such as speakers.

[0049] The hardware configuration of the learning device 1 is not limited to the configuration shown in Figure 4. For example, the learning device 1 may incorporate at least one of a display device, an input device, and a sound output device. Furthermore, the learning device 1 may be configured to include a storage device 2.

[0050] 5 is a diagram showing an example of the hardware configuration of the robot controller 3. The robot controller 3 includes, as hardware, a processor 31, a memory 32, and an interface 33. The processor 31, the memory 32, and the interface 33 are connected via a data bus 30.

[0051] The processor 31 executes a program stored in the memory 32 to function as a controller (arithmetic device) that performs overall control of the robot controller 3. The processor 31 is, for example, a processor such as a CPU, a GPU, or a TPU. The processor 31 may be composed of multiple processors.

[0052] The memory 32 is configured by various types of volatile and non-volatile memory, such as RAM, ROM, and flash memory. The memory 32 also stores programs for executing processes performed by the robot controller 3. Some of the information stored in the memory 32 may be stored in one or more external storage devices (e.g., storage device 2) that can communicate with the robot controller 3, or may be stored in a storage medium that is detachable from the robot controller 3.

[0053] The interface 33 is an interface for electrically connecting the robot controller 3 to other devices. These interfaces may be wireless interfaces such as network adapters for wirelessly transmitting and receiving data to and from other devices, or may be hardware interfaces for connecting to other devices via cables or the like.

[0054] The hardware configuration of the robot controller 3 is not limited to the configuration shown in Fig. 5. For example, the robot controller 3 may incorporate at least one of a display device, an input device, and a sound output device. The robot controller 3 may also be configured to include the storage device 2.

[0055] (3) Abstract space The robot controller 3 formulates an operation plan for the robot 5 in an abstract space based on the skill tuple. The abstract space targeted in the operation plan for the robot 5 will now be described.

[0056] FIG. 6 is a diagram showing a robot (manipulator) 5 that grasps an object and an object 6 to be grasped in real space. FIG. 7 is a diagram showing the state shown in FIG. 6 in an abstract space.

[0057] Generally, formulating a motion plan for a robot 5 performing a pick-and-place task requires rigorous calculations that take into account the shape of the end effector of the robot 5, the geometric shape of the object 6 to be grasped, the grasping position and posture of the robot 5, and the object characteristics of the object 6 to be grasped. In this embodiment, the robot controller 3 formulates a motion plan in an abstract space in which the states of each object, such as the robot 5 and the object 6 to be grasped, are abstractly (simply) represented. In the example of FIG. 7 , the abstract space defines an abstract model 5x corresponding to the end effector of the robot 5, an abstract model 6x corresponding to the object 6 to be grasped, and an executable region (see dashed-line frame 60) for the robot 5 to grasp the object 6 to be grasped. Note that, as described above, the executable state set in the abstract space is also represented as a set of combinations of initial states and target states / known task parameter values ​​in which a skill can be executed. In the example of FIG. 7, a set of combinations of initial states and target states / known task parameter values ​​in which a grasping skill can be executed is illustrated as a grasping operation executable region in a dashed frame 60. In this way, the state of the robot in the abstract space is expressed abstractly as the state of the end effector, etc. The state of each object corresponding to the operation target or environmental object is also expressed abstractly in a coordinate system based on a reference object such as a workbench.

[0058] In this embodiment, the robot controller 3 uses skills to formulate a motion plan in an abstract space that abstracts an actual system. This effectively reduces the computational cost required for motion planning, even in multi-stage tasks. In the example of Fig. 7, the robot controller 3 formulates a motion plan to execute skills for performing grasping in a graspable area (dashed frame 60) defined in the abstract space, and generates control commands for the robot 5 based on the formulated motion plan.

[0059] Hereinafter, the state of the system in real space may be represented as "x" and the state of the system in abstract space as "x'" to distinguish between them. The state x' is represented as a vector (abstract state vector). For example, in the case of a task such as pick-and-place, the abstract state vector includes a vector representing the state of the manipulated object (e.g., position, posture, velocity, etc.), a vector representing the state of the end effector of the manipulable robot 5, and a vector representing the state of environmental objects. In this way, the state x' is defined as a state vector that abstractly represents the states of some elements in the real system. Similarly, the target state / known task parameter value in the real space is defined as β g ”, and the goal state / known task parameter value in the abstract space is defined as “β g These are sometimes distinguished by writing ".

[0060] (4) Skill execution control system 8 is a diagram showing an example of the configuration of a control system related to skill execution. The processor 31 of the robot controller 3 functionally comprises a motion planning unit 34, a high-level control unit 35, and a low-level control unit 36. The system 50 corresponds to an actual system (an actual system including the robot 5). The high-level control section 35 is also called a high-level controller, and π H The high-level control unit 35 corresponds to an example of a control means. The low-level control unit 36 ​​is also called a low-level controller, and π L It is expressed as: The robot controller 3 is an example of a control device that controls the robot 5.

[0061] 8, for the sake of convenience, a balloon showing a diagram illustrating an abstract space targeted by the action planning unit 34 (see FIG. 7) is displayed in association with the action planning unit 34, and a balloon showing a diagram illustrating an actual system corresponding to the system 50 (see FIG. 6) is displayed in association with the system 50. Similarly, in FIG. 8, a balloon showing information related to an executable state set of a skill is displayed in association with the high-level control unit 35.

[0062] The motion planning unit 34 formulates a motion plan for the robot 5 based on the state x' in the abstract system and the skill database. The motion planning unit 34 expresses the target state using, for example, a logical expression based on temporal logic. The motion planning unit 34 may express the logical expression using any temporal logic such as linear temporal logic, MTL (Metric Temporal Logic), or STL (Signal Temporal Logic). The action planning unit 34 converts the generated logical formula into a sequence (action sequence) for each time step. This action sequence includes, for example, information about the skill used in each time step.

[0063] The high-level control unit 35 recognizes the skill to be executed for each time step based on the motion sequence generated by the motion planning unit 34. Then, the high-level control unit 35 determines the high-level controller "π H ", a parameter "α" to be input to the low-level control unit 36 ​​is generated.

[0064] The high-level control unit 35 generates a control parameter α as shown in the following equation (1) when the combination of the state “x0′” in the abstract space at the start of execution of the skill to be executed and the target state / known task parameter value belongs to the executable state set “χ0′” of that skill.

[0065]

number

[0066] As described above, the state at the start of skill execution is also referred to as the “initial state.” The initial state is represented, for example, as a state in an abstract space. Furthermore, if we define the approximation function of the level set function that can determine whether a skill belongs to the executable state set χ0′ as “g^”, the robot controller 3 can determine whether the state x0′ belongs to the executable state set χ0′ by determining whether equation (2) is satisfied.

[0067]

number

[0068] Equation (2) can be said to represent the constraints that determine the feasibility of a skill from a certain state. Alternatively, the approximation function g^ can be said to be a model that can evaluate whether a goal state can be reached from a certain initial state x0' under known task parameter values. The approximate function g^ is obtained by the learning device 1 through learning, as will be described later.

[0069] The set of goal states in the abstract space after the execution of the target skill is defined as χ' d " and the execution time of the target skill is denoted as "T". Also, the state when T time has elapsed since the start of skill execution is denoted as "x'(T)". By executing the skill using the low-level control unit 36, it is possible to realize equation (3).

[0070]

number

[0071] The low-level control unit 36 ​​calculates the control parameter α generated by the high-level control unit 35, the state x of the current real system obtained from the system 50, and the target state / known task parameter value β g The low-level control unit 36 ​​generates the input "u" based on the low-level controller "π L ", the input u is generated as a control command as shown in equation (4).

[0072]

number

[0073] In addition, the low-level controller π L is not limited to the above formula format, and may be a controller having various formats.

[0074] The low-level control unit 36 ​​acquires the state of the robot 5 and the environment recognized using any state recognition technology based on the measurement signals output by the measurement device 4 (which may include signals from the robot 5), etc., as state x. In FIG. 8, the system 50 is represented by a state equation shown in formula (5) using a function "f" with an input u to the robot 5 and a state x as arguments.

[0075]

number

[0076] operator" · " represents the differentiation with respect to time or the difference with respect to time.

[0077] (5) Overview of skill database updates Figure 9 is a diagram showing an example of the functional configuration of the learning device 1 regarding updating of the skill database. The processor 11 of the learning device 1 functionally comprises an abstract system model setting unit 14, a skill learning unit 15, and a skill tuple generation unit 16. Note that Figure 9 shows an example of data exchanged between each block, but is not limited to this. The same applies to the other figures.

[0078] The abstract system model setting unit 14 sets an abstract system model based on the detailed system model information. This abstract system model is a simplified version of the detailed system model specified by the detailed system model information. The detailed system model corresponds to the system 50 in FIG. 8.

[0079] The abstract system model is a model having, as a state, an abstract state vector x' that is constructed based on the state x in the detailed system model. The motion planning unit 34 formulates a motion plan using the abstract system model. The abstract system model setting unit 14 calculates the abstract system model from the detailed system model based on, for example, an algorithm stored in advance in the storage device 2 or the like.

[0080] Alternatively, information about the abstract system model may be stored in advance in the storage device 2 or the like. In this case, the abstract system model setting unit 14 may acquire information about the abstract system model from the storage device 2 or the like. The abstract system model setting unit 14 supplies information about the set abstract system model to the skill learning unit 15 and the skill tuple generation unit 16.

[0081] The skill learning unit 15 learns the control of skill execution based on the abstract system model set by the abstract system model setting unit 14, and the detailed system model information, low-level controller information, and target parameter information stored in the storage device 2. In particular, the skill learning unit 15 learns the control of skill execution based on the abstract system model set by the abstract system model setting unit 14, the detailed system model information, low-level controller information, and target parameter information stored in the storage device 2. H The low-level controller π L The skill learning unit 15 also learns the value of the control parameter α. The skill learning unit 15 also learns the level set function, and acquires training data for learning the control parameter α, for example, by using an evaluation function that evaluates the prediction accuracy of the level set function.

[0082] The skill tuple generation unit 16 generates a set of feasible states χ0′ learned by the skill learning unit 15 and a set of feasible states χ0′ learned by the high-level controller π H The skill tuple generation unit 16 generates a skill tuple containing information about the abstract system model set by the abstract system model setting unit 14, information about the low-level controller, and target parameter information. The skill tuple generation unit 16 then registers the generated skill tuple in a skill database. The data in the skill database is used by the robot controller 3 to control the robot 5.

[0083] Each component of the abstract system model setting unit 14, the skill learning unit 15, and the skill tuple generation unit 16 can be realized, for example, by the processor 11 executing a program. Alternatively, the necessary programs may be recorded in any non-volatile storage medium and installed as needed to realize each component. Note that at least a portion of these components may not necessarily be realized by software programs, but may be realized by any combination of hardware, firmware, and software. Also, at least a portion of these components may be realized using a user-programmable integrated circuit, such as an FPGA (Field-Programmable Gate Array) or a microcontroller. In this case, the integrated circuit may be used to realize a program consisting of the above components. Also, at least a portion of each component may be configured using an ASSP (Application Specific Standard Produce), an ASIC (Application Specific Integrated Circuit), or a quantum computer control chip. In this way, each component may be realized by various hardware. The same applies to other embodiments described below. Furthermore, each component may be realized by the cooperation of multiple computers, for example, using cloud computing technology.

[0084] (6) Explanation of the Skills Learning Section 10 is a diagram showing an example of the configuration of the skill learning unit 15 according to the first embodiment. Functionally, the skill learning unit 15 includes a search point set setting unit 210, a data acquisition unit 220, a prediction accuracy evaluation function learning unit 230, and a high-level controller learning unit 240.

[0085] The search point set setting unit 210 includes a search point set initialization unit 211 and a next search point set setting unit 212 . The data acquisition unit 220 includes a system model setting unit 221 , a problem setting calculation unit 222 , and a data update unit 223 . The prediction accuracy evaluation function learning unit 230 includes a level set function learning unit 231 , a prediction accuracy evaluation function setting unit 232 , and an evaluation unit 233 .

[0086] As described above, the skill learning unit 15 controls the high-level controller π H The training data for learning is generated, and the high-level controller π is generated using the generated training data. H Furthermore, the skill learning unit 15 learns the level set function. The search point set setting unit 210 is a high-level controller π H As a candidate for the task setting to be trained, we consider the initial state x s and the target state / known task parameter value β g The search point set setting unit 210 selects, from the prepared candidates, a task setting that is to be the target for acquiring training data for learning how the robot controller 3 controls the robot 5. The search point set setting unit 210 corresponds to an example of a search point setting means.

[0087] The search point set initialization unit 211 is a high-level controller π H Specifically, the search point set initialization unit 211 sets a set of candidate task settings for the initial state x s and the target state / known task parameter value β g Set a set whose elements are combinations of

[0088] The high-level controller π set by the search point set initialization unit 211 H The set of candidate task settings to be learned is called the search point set, and X ~ search The candidate task settings are also called search points. A search point is expressed as (x s ,β g ) can be expressed as Search point(x s ,β g Once the search point (xs ,β g ) can be said to indicate the behavior of the robot 5 for each task.

[0089] The next search point set setting unit 212 sets the search point set X ~ search Each element of the subset extracted by the next search point set setting unit 212 is input to the high-level controller π H It is treated as a task setting to be studied. The next search point set setting unit 212 sets the search point set X ~ search The subset extracted from X is called the search point subset. ~ check It is expressed as: search point subset X ~ check The elements of X ~ or X ~ i Here, "i" is a positive integer representing an identification number that identifies an element of the search point subset. search point subset X ~ check The elements of are also referred to as selected search points, or simply search points.

[0090] The data acquisition unit 220 acquires the search point subset X set by the next search point set setting unit 212. ~ check Element X of ~ For each, the high-level controller π H Obtain training data for learning. The system model setting unit 221 sets the search point X ~ For each experiment, a system model is set up to set up the optimal control problem.

[0091] The problem setting calculation unit 222 sets a solution search problem that indicates task execution by the robot 5, based on the settings made by the system model setting unit 221. The solution search problem here is a problem that seeks a solution that satisfies presented constraint conditions.

[0092] Specifically, the problem setting calculation unit 222 sets an optimal control problem including constraints such as constraints on the task and constraints on the robot's operation, and an evaluation function of the possibility of reaching the goal state. The optimal control problem is a problem of finding a control input that maximizes the evaluation indicated by the evaluation function value, and can be regarded as an optimization problem. In the following, we will explain an example in which a function indicating that the smaller the evaluation function value, the higher the evaluation is used as the evaluation function for the optimal control problem. In this case, when solving the optimal control problem, a solution is sought that makes the evaluation function value as small as possible, such as the minimum value of the evaluation function. However, the learning device 1 may use, as an evaluation function for the optimal control problem, a function in which the larger the function value, the higher the evaluation.

[0093] The problem setting calculation unit 222 solves the set optimal control problem and calculates a high-level controller π that minimizes the evaluation function value. H The output value of the parameter and the evaluation function value for that output value are calculated. The evaluation function value calculated by the problem setting calculation unit 222 is ~ The problem setting calculation unit 222 corresponds to an example of calculation means.

[0094] The data update unit 223 updates the data obtained by the problem setting calculation unit 222 solving the optimal control problem with the high-level controller π H and update the training data of the level set function to include it in the training data of the high-level controller π H The training data of the high-level controller π H The training data for the level set function is training data for learning the level set function. In particular, the high-level controller π H The parameter value α to be output * , high-level controller π HThe training data for learning the optimal control problem can be used as training data for the level set function. Also, the information on whether the skill can be executed or not, which is indicated by the solution of the optimal control problem, can be used as training data for the level set function. ~ j Includes:

[0095] High-level controller π H The training data is used by the robot controller 3 to H The data can be said to be training data for learning control of the robot 5 using the above. The data update unit 223 corresponds to an example of a data acquisition means. The high-level controller π handled by the data update unit 223 H The set representing the training data of D is called the acquisition data set. opt It is expressed as:

[0096] The prediction accuracy evaluation function learning unit 230 uses the acquired data set D opt The level set function and the prediction accuracy evaluation function are learned using the above, and it is determined whether or not the learning of the level set function needs to continue. As described above, the level set function is a function that indicates a set of feasible states, which is a set of combinations of states and target states / known task parameter values ​​that can reach a goal state. The prediction accuracy evaluation function is a function that indicates an evaluation of the estimation accuracy of combinations of states and target states / known task parameter values ​​that can reach a goal state using the level set function.

[0097] The training of the level set function is performed by the high-level controller π H The search point X selected for obtaining the training data ~ The problem setting calculation unit 222 calculates the high-level controller π H It is performed using data used for the training data of the above. It is considered that there is a positive correlation between the number of training data acquired by the data update unit 223 and the estimation accuracy of the level set function. The prediction accuracy evaluation function can also be said to be a function that indicates an evaluation of the acquisition status of the training data.

[0098] The level set function learning unit 231 learns the acquired data set D opt For example, the level set function learning unit 231 learns the level set function using the acquired data set D opt For each element of the equation, it is determined whether the goal state can be reached or not based on the evaluation function value calculated by the problem setting calculation unit 222. Then, the level set function learning unit 231 determines whether the goal state can be reached and the initial state x s and the target state / known task parameter value β g The combination of is used as training data to learn the level set function. The level set function learning unit 231 corresponds to an example of a level set function learning means.

[0099] The prediction accuracy evaluation function setting unit 232 learns a prediction accuracy evaluation function for the level set function learned by the level set function learning unit 231. For example, the prediction accuracy evaluation function setting unit 232 learns a prediction accuracy evaluation function for the search point X ~ , search point X ~ Based on the distribution of candidates in the space, the search point X ~ The prediction accuracy evaluation function may be learned so that a subspace with a large number of or a subspace with a high density is highly evaluated. The prediction accuracy evaluation function setting unit 232 corresponds to an example of a prediction accuracy evaluation function setting means.

[0100] The prediction accuracy evaluation function is J g^ , or J g^j Here, "j" is a positive integer representing an identification number that identifies a task. As described above, if two tasks have different unknown task parameter values, the control system 100 treats them as separate tasks.

[0101] The evaluation unit 233 uses the prediction accuracy evaluation function to evaluate the high-level controller π H The evaluation unit 233 determines whether or not it is necessary to continue acquiring training data. The evaluation unit 233 is an example of an evaluation means. High-level controller π H Whether or not it is necessary to continue acquiring training data can also be interpreted as whether or not it is necessary to continue learning the level set function. The flag indicating the determination result of the evaluation unit 233 is also referred to as a learning continuation flag.

[0102] The high-level controller learning unit 240 is configured to have the evaluation unit 233 learn the high-level controller π H If it is determined that it is not necessary to continue acquiring training data, the acquired data set D opt Using the high-level controller π H Learn about the following. For example, the high-level controller learning unit 240 may use the acquired data set D opt Among the elements of the evaluation function, an element that indicates that the evaluation function value can reach the target state is used, and the state indicated by that element is transferred to the high-level controller π H When the input to the high-level controller π is H Learn about the following. However, the high-level controller π H The learning method is not limited to a specific method.

[0103] FIG. 11 is a diagram showing an example of data input and output in the skill learning unit 15 according to the first embodiment. In the example of FIG. 11, the search point set initialization unit 211 initializes the search point set X using the target parameter information stored in the storage device 2. ~ search For example, the search point set initialization unit 211 sets the initial state x based on the target parameter information. si and the goal state / known task parameter value β g All possible combinations of ~ search It may be set as an element of Search point set X by the search point set initialization unit 211 ~ search The setting of the search point set X ~ search This corresponds to the initial setting of the search point set X ~ search is updated by the next search point set setting unit 212.

[0104] The next search point set setting unit 212 sets the search point set X ~ search From the search point subset X ~ check Specifically, the next search point set setting unit 212 extracts the search point set X ~ search Read one or more elements from the search point subset X ~ check Then, the next search point set setting unit 212 sets the read elements as elements of the search point subset X ~ check The elements set to are added to the search point set X ~ search Remove from the elements of .

[0105] When the prediction accuracy evaluation function setting unit 232 has learned the prediction accuracy evaluation function, the next search point set setting unit 212 sets the search point subset X ~ check In particular, the next search point set setting unit 212 sets the search point set X ~ search Among the elements of , the elements whose prediction accuracy evaluation function value indicates that the estimation accuracy of the level set function is lower than the predetermined condition are included in the search point subset X ~ check Set it as an element of .

[0106] The method for determining whether the estimation accuracy is lower than the predetermined condition is not limited to a specific method. For example, if a larger prediction accuracy evaluation function value indicates lower accuracy, the estimation accuracy being lower than the predetermined condition may be, but is not limited to, that the prediction accuracy evaluation function value is greater than a predetermined threshold.

[0107] The system model setting unit 221 sets the search point subset X ~ checkFor each element of the above, various settings are made to set the optimal control problem. For example, the system model setting unit 221 sets the low-level controller π based on the detailed system model information, low-level controller information, and target parameter information stored in the storage device 2, and the abstract system model set by the abstract system model setting unit 14. l A system model, constraints on the parameters of the system model, and an evaluation function for the possibility of reaching the goal state are set.

[0108] The system model here refers to a model of the target system, such as a motion model of the target system. Constraints on the parameters of the system model are constraints on the values ​​that the parameters of the system model can take, such as constraints on the specifications of the devices included in the target system and physical constraints. The system model and the constraints on the parameters of the system model are used as part of the constraints in the optimal control problem handled by the problem setting calculation unit 222.

[0109] The system model setting unit 221 sets the low-level controller π l , a system model, parameters of the system model, an evaluation function for the possibility of reaching the goal state, and a search point X ~ i Then, information regarding the time limit for skill execution, such as execution time T, is output to the problem setting calculation unit 222.

[0110] The problem setting calculation unit 222 calculates the search point X based on the information from the system model setting unit 221. ~ i An optimal control problem is set for each parameter, and a solution to the set optimal control problem is searched for. As described above, the optimal control problem is, for example, a problem of determining a control input that minimizes the value of an evaluation function. Specifically, the optimal control problem here is a problem of determining a control input that minimizes the value of an evaluation function under constraints imposed by the operating environment, etc., when an initial state and an evaluation function are given. The problem setting calculation unit 222 sets an evaluation function for the possibility of reaching the target state as an evaluation function in the optimal control problem, and sets other various settings as constraint conditions in the optimal control problem.

[0111] The problem setting calculation unit 222 calculates a high-level controller π that minimizes the evaluation function value under the constraints in the optimal control problem. H The problem setting calculation unit 222 calculates the output value of the search point X ~ i and the high-level controller π that minimizes the evaluation function value. H Output value α * i And the evaluation function value g * i Combination with (X ~ i ,g * i ,α * i ) to the data update unit 223.

[0112] For example, the problem setting calculation unit 222 may use, as the evaluation function for the optimal control problem, an evaluation function g such that the state x' is the target state when the formula (6) holds.

[0113]

number

[0114] If equation (6) holds, then state x' is the target state, as expressed by equation (7).

[0115]

number

[0116] x d ' represents the target state set. If we represent the mapping from the state x of the detailed system model to the state x' of the abstract system model by γ, we can obtain equation (8) from equation (7).

[0117]

number

[0118] Minimizing the value of the evaluation function g in the optimal control problem is expressed as in equation (9).

[0119]

number

[0120] As mentioned above, T represents the time required to perform the skill. g(γ(x(T)),β g ) represents the evaluation function value for the state x(T) at the end of the skill. If this evaluation function value is 0 or less, it can be determined that the goal state can be reached by executing the skill. As mentioned above, α is the high-level controller π H Equation (9) represents the output of the high-level controller π that minimizes the value of the evaluation function g. H This represents the determination of the output α. The system model in the optimal control problem can be expressed as in equation (10).

[0121]

number

[0122] As mentioned above, τ j represents the unknown task parameters. The time t is expressed as in equation (11).

[0123]

number

[0124] The inequality constraints in the optimal control problem can be expressed as in equation (12).

[0125]

number

[0126] c is a function representing a constraint condition, and is set based on target parameter information, for example. The state at time 0 is the initial state and is expressed as equation (13).

[0127]

number

[0128] The fact that γ is a mapping from the state x of the detailed system model to the state x' of the abstract system model can be expressed as in equation (14).

[0129]

number

[0130] The problem setting calculation unit 222 calculates the output α of the high-level controller so as to minimize the value of the evaluation function g as shown in equation (9) under the constraints of, for example, equations (10) to (14). * , and the value of the evaluation function g at that time g * As shown in equation (6), g * If ≦0, then from the initial state at this time, the output of the high-level controller is * By executing the skill as above, it can be determined that the goal state can be reached.

[0131] The problem setting calculation unit 222 calculates the minimum value g of the obtained evaluation function. * , and the output α of the high-level controller at that time * the initial state x s , and the target state / known task parameter value β g , and outputs it to the data update unit 223. Alternatively, the problem setting calculation unit 222 outputs the output α *In addition to or instead of this, information indicating whether the goal state can be reached may be output to the data update unit 223. The data update unit 223 updates this data with the high-level controller π H Include it in the training data used to learn.

[0132] The method by which the problem setting calculation unit 222 solves the optimal control problem is not limited to a specific method. For example, the problem setting calculation unit 222 may use a known algorithm as a solution search algorithm for optimal control problems, or a known algorithm as a round-robin search problem for optimization problems. Alternatively, the problem setting calculation unit 222 may simulate the behavior of the robot 5 and learn behaviors such that the evaluation function value is as small as possible, such as by reinforcement learning.

[0133] For example, when the function f in equation (10) is analytically obtained, the problem setting calculation unit 222 can solve the optimal control problem using any optimal control algorithm such as the Direct Collocation method or Differential Dynamic Programming (DDP).

[0134] On the other hand, when the function f cannot be analytically obtained, for example, when a simulator is used as the function f, the problem setting calculation unit 222 can solve the optimal control problem using a black-box optimization method such as path integral control or a model-free optimal control method. In this case, the problem setting calculation unit 222 determines the control parameter α according to the problem of minimizing the evaluation function g, based on the function c representing the constraint conditions.

[0135] Here, when generating the skill of the grasping operation in the pick-and-place task shown in FIG. 6, the target parameter information and the low-level controller π L A specific example of this will be described. Here, "generating a skill" means learning a skill for a task different from the task for which the skill has already been learned. As mentioned above, a different task is a task with different values ​​for unknown task parameters.

[0136] Here, a physical simulator based on the state x, the input u to the robot 5, and the contact force F that is the force for gripping the object to be grasped 6 is used as the system model shown in equation (10). In this case, the equation for determining whether the goal state can be reached is expressed as equation (15).

[0137]

number

[0138] If equation (15) holds, it can be determined that the goal state can be reached. In addition, the execution time information of the target parameter information includes the upper limit of the skill execution time T, "T max ” (T≦T max ) is included. The general constraint condition information of the target parameter information is also included information representing the constraint equations for the state x, input u, and contact force F as shown in Equation (16).

[0139]

number

[0140] For example, this constraint is the upper limit of the contact force F, "F max ” (F≦F max ), the limit of the range of motion (or speed) "x max ” (|x|≦x max ), the upper limit of input u is "u max ” (|u|≦u max ) is a comprehensive expression.

[0141] Also, the low-level controller π L is assumed to be a servo controller using PID, for example. Here, the state of the robot 5 is expressed as "xr ”, and the target trajectory of the state of robot 5 is “x rd ", then the input u is expressed as, for example, equation (17).

[0142]

number

[0143] target trajectory x rd is expressed as, for example, equation (18).

[0144]

number

[0145] In equations (17) and (18), the control parameters based on the output α of the high-level controller πH are the coefficients of the target trajectory polynomial and the gain of the PID control, and are expressed as in equation (19).

[0146]

number

[0147] The problem setting calculation unit 222 solves the optimal control problem and obtains the optimal value (α * ) is calculated.

[0148] The data update unit 223 updates the data output from the problem setting calculation unit 222 (X ~ i ,g * i ,α * i ) to obtain the data set D opt The acquired data set D opt Update.

[0149] As described above, the level set function learning unit 231 learns the acquired data set D optThe level set function learning unit 231 outputs the obtained level set function to the prediction accuracy evaluation function setting unit 232.

[0150] For example, the level set function learning unit 231 uses the acquired data set D opt The evaluation function value shown in is compared with a predetermined threshold, and the acquired data set D opt In the example of equations (8) and (9), the level set function learning unit 231 determines whether the target state can be reached from the initial state shown in * Whether or not the goal state can be reached is determined based on whether or not is less than or equal to 0.

[0151] Then, the level set function learning unit 231 learns the acquired data set D opt The level set function is learned using a combination of the state shown in the table, the goal state, and the result of determining whether the goal state can be reached as training data.

[0152] Here, the initial state x0' in the abstract state and the goal state / known task parameter value β g The optimal value g of the evaluation function g for * A function that outputs g * (x0',β g ) The executable state set χ0' of the target skill is expressed as in equation (20).

[0153]

number

[0154] The level set function learning unit 231 learns the acquired data set D opt The initial state x0' and the goal state / known task parameter value β g ' and the function value g *Based on a plurality of pairs of χ and χ, the level set function learning unit 231 learns a level set function that represents the feasible state set χ'. For example, the level set function learning unit 231 calculates the level set function using a level set estimation method, which is an estimation method using Gaussian process regression based on the idea of ​​Bayesian optimization. Here, this level set function is called g GP It is expressed as:

[0155] Note that this level set function g GP may be defined using the mean function of a Gaussian process obtained through the level set estimation method, or may be defined as a combination of the mean function and the variance function. The method by which the level set function learning unit 231 learns the function indicating the feasible state set is not limited to a specific method. For example, the level set function learning unit 231 may obtain the level set function using Truncated Variance Reduction (TruVaR), which is an estimation method using Gaussian process regression similar to the level set estimation method.

[0156] As described above, the level set function may be any model that evaluates the initial state that can be reached with respect to the desired state. H Output value α * is the initial state x0' and the goal state / known task parameter value β g ' and the evaluation function value g * It can be said that the level set function is determined based on the set of π and π. By determining the level set function, it is possible to evaluate the reachable states and known task parameter values, and therefore it is possible to determine the control parameters that can achieve the desired state of the system. Here, the high-level controller π H Output value α * are examples of control parameters.

[0157] Furthermore, a control device for a robot or the like may use a level set function to determine whether a desired state can be reached from an initial state under given known task parameter values, and if the control device determines that a desired state can be reached, the control device may control the controlled object, such as a robot, using control parameters corresponding to the initial state.

[0158] In order to reduce the calculation cost of the level set function, the level set function learning unit 231 may acquire a simplified level set function by polynomial approximation or the like through learning. In this case, the level set function is represented by g^. g^ is also called a level set approximation function. The level set function learning unit 231 may learn a level set approximation function g^ that satisfies equation (21).

[0159]

number

[0160] As described above, the prediction accuracy evaluation function setting unit 232 sets a prediction accuracy evaluation function that indicates an evaluation of the level set function learned by the level set function learning unit 231. The prediction accuracy evaluation function setting unit 232 outputs the obtained prediction accuracy evaluation function to the evaluation unit 233.

[0161] For example, the prediction accuracy evaluation function setting unit 232 sets the search point X ~ , search point X ~ A function indicating an evaluation of the distribution of the search points X in the space of candidates may be learned as a prediction accuracy evaluation function. ~ The space of candidates for the search point X ~ The prediction accuracy evaluation function setting unit 232 determines the search point X ~ The space that the domain of ~ Alternatively, the search point X ~ The space of candidates is the search point set X ~ searchIt may be the initial value of

[0162] For example, as a prediction accuracy evaluation function, search point X ~ The candidate of X is used as an argument, and the search point X ~ Alternatively, a function may be used that outputs, as a function value, an evaluation value for the possibility of reaching the goal state indicated by the level set function for each candidate. Then, the prediction accuracy evaluation function setting unit 232 sets the search point X ~ A trained search point X within a given distance from the candidate ~ The prediction accuracy evaluation function value may be calculated so that a larger number of values ​​indicates a higher evaluation. Alternatively, as will be described in the third embodiment, when the variance of the level set function values ​​is obtained, the prediction accuracy evaluation function setting unit 232 may set the prediction accuracy evaluation function so that the smaller the variance of the level set function values, the higher the evaluation. However, the method by which the prediction accuracy evaluation function setting unit 232 learns the prediction accuracy evaluation function is not limited to a specific method. In the following, unless there is a need to distinguish between them, we use the level set function g GP and the level set function ĝ are collectively referred to as the level set function ĝ.

[0163] As described above, the evaluation unit 233 uses the prediction accuracy evaluation function to determine the high-level controller π H The evaluation unit 233 determines whether or not it is necessary to continue acquiring the training data. The evaluation unit 233 sets the result of the determination in a learning continuation flag. For example, the evaluation unit 233 evaluates the search point X ~ The minimum value of the prediction accuracy evaluation function in the space of candidates may be calculated. The minimum value of the prediction accuracy evaluation function here refers to the lowest evaluation value. If the minimum value of the prediction accuracy evaluation function is evaluated lower than a predetermined threshold, the evaluation unit 233 may determine that it is necessary to continue acquiring training data. On the other hand, if the minimum value of the prediction accuracy evaluation function is evaluated higher than or equal to a predetermined threshold, the evaluation unit 233 may determine that it is not necessary to continue acquiring training data.

[0164] Alternatively, the evaluation unit 233 may ~ Alternatively, prediction accuracy evaluation function values ​​may be sampled in the space of candidates, and the necessity of continuing to acquire training data may be determined based on the lowest evaluation value among the obtained prediction accuracy evaluation function values. However, the evaluation unit 233 is H The method of determining whether or not it is necessary to continue acquiring training data is not limited to a specific method. For example, the evaluation unit 233 may determine whether or not it is necessary to continue acquiring training data based on a predetermined learning condition in addition to the value of the prediction accuracy evaluation function. The learning condition here may be various conditions. For example, when the number of times training data has been acquired reaches a predetermined number or more, the evaluation unit 233 may determine that it is not necessary to continue acquiring training data even if the evaluation indicated by the prediction accuracy evaluation function does not reach the predetermined evaluation.

[0165] As described above, the high-level controller learning unit 240 is configured to determine whether the evaluation unit 233 is a high-level controller π H If it is determined that it is not necessary to continue acquiring training data, the acquired data set D opt Using the high-level controller π H Learn about the following. Specifically, the high-level controller learning unit 240 uses the acquired data set D opt Among the elements of , the elements that can reach the target state are controlled by the high-level controller π H The initial state x0' and the goal state / known task parameter value β g For the input of ', the output value α * The high-level controller π H Learn about the following.

[0166] The high-level controller learning unit 240 learns the high-level controller π HThe model used for learning the above can be various models, for example, a neural network, a Gaussian process regression, or a support vector regression, but is not limited to these.

[0167] (7) Processing flow 12 is a diagram showing an example of a skill database update process performed by the learning device 1 according to the first embodiment. The learning device 1 executes the process of FIG. 12 for each skill to be generated.

[0168] (Step S101) The search point set initialization unit 211 initializes the search point set X ~ search and the acquired data set D opt Perform initial settings. For example, the search point set initialization unit 211 initializes the initial state x s and the goal state / known task parameter value β included in the goal state / known task parameter information. g Each of the arbitrary combinations of X and X is a search point set ~ search As an element of the search point set X ~ search Generate. The search point set initialization unit 211 also initializes the acquired data set D opt Set the value of to the empty set. After step S101, the process proceeds to step S102.

[0169] (Step S102) The next search point set setting unit 212 sets the search point set X ~ search Specifically, the next search point set setting unit 212 extracts a subset from the search point set X ~ search A subset of X is a search point subset X ~ check Then, the next search point set setting unit 212 sets the set search point subset X ~ checkEach element of the search point set X ~ search Exclude from. As shown in equation (22), the search point subset X ~ check is the initial state x si and the target state / known task parameter value β gi It has elements of combination with.

[0170]

number

[0171] The next search point set setting unit 212 sets the set subset X ~ check Each element of the search point set X ~ search The process of excluding from can be expressed as in equation (23).

[0172]

number

[0173] Here, "-" indicates the removal of a subset from a set. After step S102, the process proceeds to step S103.

[0174] (Step S103) The learning device 1 selects a subset X of the search point set. ~ check Search point X, which is an element of ~ In the loop L11, the number of times the loop is repeated is represented by "i". Also, the search point X ~ The target search point X ~ i It is also called. After step S103, the process proceeds to step S104.

[0175] (Step S104) The system model setting unit 221 sets the target search point X ~i For example, the system model setting unit 221 performs various settings for setting the optimal control problem based on the low-level controller π l A system model, constraints on the parameters of the system model, and an evaluation function for the possibility of reaching the goal state are set. After step S104, the process proceeds to step S105.

[0176] (Step S105) The problem setting calculation unit 222 sets an optimal control problem based on the setting by the system model setting unit 221 in step S104. Then, the problem setting calculation unit 222 solves the set optimal control problem and calculates the output α of the high-level controller so that the evaluation function value becomes as small as possible. * , and the value of the evaluation function g at that time g * is obtained as the solution. After step S105, the process proceeds to step S106.

[0177] (Step S106) The data update unit 223 updates the acquired data set D opt Specifically, the data updating unit 223 updates the subset X of the search point set. ~ check the i-th element of X ~ i and the result of the task success or failure g * i and the obtained control parameter α * i Combination with (X ~ i ,g * i ,α * i ) to obtain the data set D opt Add it as an element of . The data update unit 223 acquires the data set D opt The process of updating is expressed as in equation (24).

[0178]

number

[0179] "{(X ~ i ,g * i ,α * i )}" is (X ~ i ,g * i ,α * i ) as an element, and represents a set of one element. After step S106, the process proceeds to step S107.

[0180] (Step S107) The learning device 1 performs the termination process of the loop L11. Specifically, the learning device 1 performs the termination process of the subset X ~ check It is determined whether the processing of loop L11 has been performed for all elements in the learning device 1. If it is determined that there are elements for which the processing of loop L11 has not yet been performed, the learning device 1 continues to perform the processing of loop L11 for the elements for which the processing of loop L11 has not yet been performed. In this case, the processing returns to step S103. On the other hand, a subset X of the search point set ~ check If it is determined that the processing of loop L11 has been performed for all elements of, the learning device 1 ends loop L11. In this case, the processing proceeds to step S111.

[0181] (Step S111) The level set function learning unit 231 learns the acquired data set D opt The level set function g^ is learned based on this. After step S111, the process proceeds to step S112.

[0182] (Step S112) The prediction accuracy evaluation function setting unit 232 determines the prediction accuracy evaluation function J based on the level set function g. g^ Set. After step S112, the process proceeds to step S110.

[0183] (Step S113) The evaluation unit 233 calculates the prediction accuracy evaluation function J g Based on this, the evaluation unit 233 determines whether or not it is necessary to continue learning the level set function ĝ. g In addition, it may be determined whether or not the learning of the level set function g^ needs to continue based on a predetermined learning condition. If the evaluation unit 233 determines that it is necessary to continue learning the level set function g^ (step S113: YES), the process proceeds to step S121. On the other hand, if the evaluation unit 233 determines that it is not necessary to continue learning the level set function g^ (step S113: NO), the process proceeds to step S131.

[0184] (Step S121) The next search point set setting unit 212 calculates the prediction accuracy evaluation function J g Based on this, the search point set X ~ search Subset X from ~ check Specifically, the next search point set setting unit 212 extracts the prediction accuracy evaluation function J g Based on this, the search point set X ~ search A subset X of ~ check Then, the next search point set setting unit 212 sets the set subset X ~ check Each element of the search point set X ~ search Exclude from. After step S121, the process returns to step S103.

[0185] (Step S131) The high-level controller learning unit 240 learns the acquired data set D opt Using the high-level controller π H Learn about the following. After step S131, the learning device 1 ends the processing of FIG.

[0186] As described above, the search point set setting unit 210 sets the search points (x s ,β g ) search point X, which is the target of acquiring training data for learning the control of robot 5. ~ Select . The problem setting calculation unit 222 calculates the selected search point X ~ The information indicating whether the action indicated by is executable or not and the selected search point X ~ A high-level controller π controls the robot 5 in response to the motion indicated by H and the output value to be output. The data update unit 223 updates the selected search point X ~ and the selected search point X ~ The information indicating whether the action indicated by is executable or not and the selected search point X ~ For the operation shown by H and the output value to be output by the high-level controller π H The training data is acquired for learning the control of the robot 5 by the The evaluation unit 233 determines whether or not to continue acquiring training data based on the evaluation of the acquisition status of the training data.

[0187] According to the learning device 1, it is possible to determine whether or not it is necessary to continue learning the control of the robot 5, and it is possible to eliminate unnecessary learning, thereby making learning efficient.

[0188] In addition, the level set function learning unit 231 s ,β g ) and the search point (x s ,β g The problem setting calculation unit 222 learns the level set function ĝ, which outputs an estimate of whether the action indicated by the search point (x s ,β g ) is based on the evaluation result of whether the action indicated by the

[0189] The prediction accuracy evaluation function setting unit 232 sets the search point (x s ,β g) and the search point (x s ,β g ) is a prediction accuracy evaluation function J that outputs an evaluation value of the estimation accuracy of the level set function g^. g^ The evaluation unit 233 sets the prediction accuracy evaluation function J g^ Based on this, it is determined whether to continue acquiring training data.

[0190] According to the learning device 1, it is possible to determine whether to continue acquiring training data using the level set function g^. The level set function g^ is used to select skills when the robot controller 3 controls the robot 5. According to the learning device 1, the amount of work required just to determine whether to continue acquiring training data is relatively small, and in this respect, it is possible to efficiently determine whether to continue acquiring training data.

[0191] The search point set setting unit 210 also uses the prediction accuracy evaluation function J g^ The evaluation value of the search point (x s ,β g ) is selected as a target for acquiring training data for controlling the robot 5. As a result, in the learning device 1, the high-level controller π H It is possible to obtain training data that show inputs and outputs that are likely to be inaccurate, and to use the high-level controller π H This allows for efficient learning.

[0192] Also, the search point (x s ,β g ) includes known task parameters, which are parameter values ​​of the skills in which the operation of the controlled object is modularized. As a result, the learning device 1 can express differences in the movements of the robot 5 that can be expressed by parameter values ​​using skill parameter values, and can perform control learning so as to apply the same skill to different movements.

[0193] Also, the search point (x s ,β g) is composed of a combination of the initial state of the robot 5 and its operating environment at the start of the skill, the known parameter values ​​of the skill, and the target state of the robot 5 and its operating environment at the end of the skill. As a result, the learning device 1 can calculate the high-level controller π H It is possible to learn the high-level controller π H and low-level controller π L This makes it possible to perform learning more efficiently than when learning control equivalent to both of the above.

[0194] The robot controller 3 also uses the high-level controller π obtained by learning using the training data acquired by the learning device 1. H Equipped with. According to the robot controller 3, when learning the robot controller 3, it is possible to determine whether or not it is necessary to continue learning the control of the robot 5, and learning can be carried out efficiently in that unnecessary learning can be eliminated.

[0195] The robot controller 3 also includes a high-level controller π that controls the robot 5 in accordance with the size of the object to be grasped so that the robot 5 can grasp each of the objects to be grasped that have different sizes. H Equipped with. The robot controller 3 is expected to be able to control the robot 5 with high precision according to the size of the object to be grasped. Second Embodiment When the data acquisition unit 220 acquires data, the high-level controller learning unit 240 learns the high-level controller π H The second embodiment will be described in detail below. The configuration of the control system 100 in the second embodiment is the same as that in the first embodiment. The second embodiment will also be described using the configuration of the control system 100 shown in Figs.

[0196] 13 is a diagram showing an example of data input / output in the skill learning unit 15 according to the second embodiment. In the second embodiment, when the data acquisition unit 220 acquires data, the high-level controller learning unit 240 learns the high-level controller, and the high-level controller π obtained by the learning * H to the data acquisition unit 220. Outputting the high-level controller can be performed by outputting the setting values ​​of the parameters of a predictor that constitutes the high-level controller, such as a neural network or a Gaussian process. In other respects, the input and output of data shown in FIG. 13 is similar to the input and output of data in the first embodiment described with reference to FIG.

[0197] 14 is a diagram showing an example of a skill database update process performed by the learning device 1 according to the second embodiment. The learning device 1 executes the process of FIG. 14 for each skill to be generated. Steps S201 to S204 in FIG. 14 are the same as steps S101 to S104 in FIG. The loop from steps S203 to S207 in FIG. 14 is denoted as loop L21.

[0198] (Step S205) The problem setting calculation unit 222 sets an optimal control problem, solves the set optimal control problem, and calculates a high-level controller π that minimizes the evaluation function value, as described in step S105. H The output and the evaluation function value at that time are calculated.

[0199] On the other hand, in step S205, the problem setting calculation unit 222 sets the high-level controller π H If there is a high-level controller π H In order to avoid deviation from the output value due to the high level controller π H For example, the problem setting calculation unit 222 calculates the output of the obtained high-level controller π H and the high-level controller π HThe error norm term with respect to the output value of the high-level controller π may be included in the evaluation function of the optimal control problem. Then, the problem setting calculation unit 222 may obtain a solution to the optimal control problem so that the evaluation function value is as small as possible. In this way, the problem setting calculation unit 222 makes the original evaluation function value as small as possible and also minimizes the value of the high-level controller π H The output value of the high-level controller π H A solution close to the output value of can be obtained.

[0200] Steps S206 and S207 are similar to steps S106 and S107 in FIG. In step S207, if the learning device 1 ends the loop L21, the process proceeds to step S211.

[0201] (Step S211) The high-level controller learning unit 240 learns the high-level controller π H The criterion for determining whether or not it is necessary to continue learning the high-level controller π is not limited to a specific one. For example, the high-level controller learning unit 240 may determine whether or not it is necessary to continue learning the high-level controller π obtained by solving the optimal control problem in step S205. H and the output of the high-level controller π H When the difference between the output obtained using H It may be determined that there is no need to continue learning.

[0202] In step S211, the high-level controller π H If the high-level controller learning unit 240 determines that learning needs to be continued (step S211: YES), the process proceeds to step S221. On the other hand, the high-level controller π H If the high-level controller learning unit 240 determines that it is not necessary to continue learning (step S211: NO), the process proceeds to step S231.

[0203] (Step S221) The high-level controller learning unit 240 learns the acquired data set D optUsing the high-level controller π H The method in which the high-level controller learning unit 240 learns the high-level controller in step S221 is the same as in step S131 in FIG. 12. In step S221, the acquired data set D opt The difference from step S131 is that the is being generated. After step S221, the process returns to step S203.

[0204] Steps S231 to S233 are the same as steps S111 to S113 in FIG. In step S233, if the evaluation unit 233 determines that it is necessary to continue learning the level set function ĝ (step S233: YES), the process proceeds to step S241. On the other hand, if the evaluation unit 233 determines that it is not necessary to continue learning the level set function ĝ (step S233: NO), the process proceeds to step S251.

[0205] Step S241 is the same as step S121 in Fig. 12. After step S241, the process returns to step S203. Step S251 is the same as step 131 in Fig. 12. After step S251, the learning device 1 ends the processing in Fig. 14.

[0206] Third Embodiment In the third embodiment, an example will be described in which the learning device 1 learns skills in response to differences in tasks that are difficult to express using parameter values. Specifically, in addition to the learning in the first embodiment, the learning device 1 learns metaparameter values ​​of the predictors that constitute the level set function and the predictors that constitute the high-level controller. When acquiring training data for a new task and learning skills for performing the task, the learning device 1 learns and sets metaparameter values ​​in advance using training data that has already been acquired so that the prediction accuracy of these predictors is as high as possible. The learning device 1 may perform the learning according to the third embodiment in addition to the learning according to the second embodiment. That is, it is also possible to implement the second embodiment and the third embodiment in combination.

[0207] In the third embodiment, it is assumed that tasks are generated according to some probability distribution, and that correct answer data for input and output of a predictor follows some probability distribution determined for each task. Tasks are generated according to some probability distribution. j τ can be expressed as τ, where τ represents the probability distribution that the task follows. j represents a task. The fact that the correct data of the input and output of the predictor follows some probability distribution determined for each task is called S j ~D j It can be expressed as: D j is the task τ j The probability distribution is determined by S j is the task τ j This represents the correct data for the input and output of the predictor in this case.

[0208] Fig. 15 is a diagram showing an example of the configuration of the skill learning unit 15 according to the third embodiment. In the configuration shown in Fig. 15, the skill learning unit 15 includes a search task setting unit 250 and a meta parameter processing unit 260 in addition to the units shown in Fig. 10. In other respects, the configuration of the control system in the third embodiment is the same as that in the first embodiment. The third embodiment will also be described using the configuration of the control system 100 shown in Figures 1 to 9.

[0209] The search task setting unit 250 sets a task to be learned by the learning device 1. A task to be learned by the learning device 1, which is set by the search task setting unit 250, is also referred to as a search task. The search task setting unit 250 assumes a probability distribution T that the generated task will follow, and sets the search task based on the assumed probability distribution T. The method by which the search task setting unit 250 assumes the probability distribution T that the generated task will follow is not limited to a specific method. For example, the probability distribution T may be set in advance, but is not limited to this.

[0210] The meta-parameter processing unit 260 includes a predictor that configures a level set function and a high-level controller π H The meta parameter values ​​of the predictors that make up the predictor are learned, and the meta parameter values ​​obtained by learning are set in these predictors. In the third embodiment, a predictor constituting a level set function and a high-level controller π H The predictor used for the prediction is a learning model predictor in which parameter values ​​are set according to a probability distribution, such as a Bayesian neural network or Gaussian process. The meta parameter processing unit 260 learns and sets the probability distribution to which these parameter values ​​follow as meta parameter values. Furthermore, the meta parameter processing unit 260 evaluates the prediction accuracy of the predictor for which the meta parameters have been set, and determines whether or not to continue learning the meta parameter values ​​based on the evaluation results.

[0211] Fig. 16 is a diagram showing an example of data input / output in the skill learning unit 15 according to the third embodiment. As explained with reference to Fig. 15, in the configuration shown in Fig. 16, the skill learning unit 15 also includes a search task setting unit 250 and a meta parameter processing unit 260 in addition to the units shown in Fig. 11.

[0212] The search task setting unit 250 sets a search task in response to input task parameter information. The task parameter information includes information related to the probability distribution T of the task to be generated. For example, the task parameter information may be information indicating the probability distribution T of the task to be generated, and the search task setting unit 250 may set a search task in accordance with this probability distribution T.

[0213] The search task setting unit 250 repeats setting of the search task while the learning continuation flag for the unknown task parameter indicates continuation of learning. The learning continuation flag for the unknown task parameter is a flag that indicates whether to continue learning of the meta parameter value of the predictor. While the learning continuation flag for the unknown task parameter indicates continuation of learning, the search task setting unit 250 sets the next search task each time the learning device 1 finishes learning related to the search task.

[0214] In the third embodiment, the learning continuation flag set by the evaluation unit 233 is also referred to as a learning continuation flag for a known task parameter in order to distinguish it from a learning continuation flag for an unknown task parameter. j " or "j" may be used to indicate that the data is for each task.

[0215] Each time the search task setting unit 250 sets a search task, the learning device 1 sets a task τ j In particular, the search point set initialization unit 211 initializes the search point set X ~ search Furthermore, the system model setting unit 221 performs various settings for setting the optimal control problem according to the search task.

[0216] The meta parameter processing unit 260 processes the entire acquired data set D optall The above-mentioned meta parameter value learning and the decision on whether to continue learning the meta parameter value are performed using the whole acquired data set D optall is the total acquired data set D acquired by the data update unit 223. opt,j It is a combination of the above.

[0217] For example, the data update unit 223 updates the entire acquired data set D optall The initial value of is set to 0, and the acquired data set D opt,j Each time a data set is generated, the acquired data set D opt,j The total acquired data set D optall may be coupled to Acquired data set D opt,j The total acquired data set D optall The process of combining can be expressed as in equation (25).

[0218]

number

[0219] The meta parameter values ​​learned by the meta parameter processing unit 260 are set in the predictors that constitute the level set set and the predictors that constitute the high-level controller. Also, as described above, while the learning continuation flag for the unknown task parameter set by the meta parameter processing unit 260 indicates that learning is continuing, the search task setting unit 250 sets the next search task each time the learning device 1 finishes learning related to a search task.

[0220] 17 is a diagram showing an example of the configuration of the meta parameter processing unit 260. In the configuration shown in FIG.

[0221] The meta parameter processing unit 260 includes a meta parameter individual processing unit 261 for each predictor to be learned. In the example of FIG. 16, the level set function and the high-level controller π H These are configured using a predictor and are the targets of learning meta parameter values. In this case, the meta parameter processing unit 260 includes two meta parameter individual processing units 261.

[0222] However, the number of meta parameter individual processors 261 included in the meta parameter processor 260 is not limited to two. For example, the level set function and the high-level controller π H In addition, there may be other objects configured using a predictor and subject to learning of meta parameter values. In this case, the meta parameter processing unit 260 may be provided with an individual meta parameter processing unit 261 for each object configured using a predictor and subject to learning of meta parameter values.

[0223] When distinguishing between the individual meta parameter individual processing units 261, they are referred to as meta parameter individual processing unit 261-1, meta parameter individual processing unit 261-2, ..., meta parameter individual processing unit 261-N, where N is a positive integer indicating the number of meta parameter individual processing units 261 included in the meta parameter processing unit 260.

[0224] The meta parameter individual processing unit 261 learns the meta parameter values ​​of the predictor. If the predictor has multiple meta parameters, the meta parameter individual processing unit 261 learns the value of each meta parameter. For example, if each predictor is configured using a Bayesian neural network and has weight coefficients between nodes and biases for each node as parameters, the probability distributions that each of these parameters follows correspond to meta parameters. The meta parameter individual processing unit 261 learns the values ​​of each of these meta parameters.

[0225] Furthermore, the meta parameter individual processing unit 261 sets the value of a learning continuation flag for each predictor, which indicates whether or not learning of the meta parameter value needs to be continued for the target predictor. The learning continuation flag for each predictor is also referred to as an individual learning continuation flag. The learning continuation flag integrating unit 262 integrates the values ​​of the individual learning continuation flags to set the value of the learning continuation flag for the unknown task parameter. The learning continuation flag integrating unit 262 corresponds to an example of a learning continuation determination integration means.

[0226] FIG. 18 is a diagram showing an example of data input / output in the meta parameter processing unit 260. As shown in FIG. As described above, the meta parameter individual processing unit 261 is provided for each predictor targeted by the meta parameter processing unit 260. The meta parameter individual processing unit 261 processes all the acquired data D optall and the meta-learning execution flag or the internal learning evaluation value, the meta parameter individual processing unit 261 outputs the value of the target meta parameter and also sets the value of the individual learning continuation flag.

[0227] The meta-learning execution flag is a flag that indicates whether or not to perform learning of meta-parameter values. For example, the entire acquired data set D optall When a predetermined number of pieces of data (set elements) for each task are accumulated in the meta parameter processing unit 260, the data update unit 223 may set the value of the meta-learning execution flag to a value indicating that learning of meta parameter values ​​will be performed. Also, when the meta parameter processing unit 260 ends learning of meta parameter values, the data update unit 223 may set the value of the meta-learning execution flag to a value indicating that learning of meta parameter values ​​will not be performed.

[0228] The internal learning evaluation value is a value indicating an evaluation of the prediction accuracy of a predictor. For example, when learning of meta parameter values ​​starts, the meta parameter individual processing unit 261 may calculate a generalization error of the meta parameter. Then, the meta parameter processing unit 260 may calculate an internal learning evaluation value indicating a comprehensive evaluation of all predictors that are the learning targets of meta parameter values, based on the generalization error of the meta parameter.

[0229] The learning continuation flag integrating unit 262 integrates the values ​​of the individual learning continuation flags to set the value of the learning continuation flag for the unknown task parameter. For example, if the values ​​of one or more individual learning continuation flags indicate that learning continuation is necessary, the learning continuation flag integrating unit 262 sets the value of the learning continuation flag for the unknown task parameter to a value indicating that learning continuation is necessary. On the other hand, if the values ​​of all the individual learning continuation flags indicate that learning continuation is not necessary, the learning continuation flag integrating unit 262 sets the value of the learning continuation flag for the unknown task parameter to a value indicating that learning continuation is not necessary.

[0230] Fig. 19 is a diagram showing a first example of the configuration of the meta parameter individual processing unit 261. In the configuration shown in Fig. 19, the meta parameter individual processing unit 261 includes a training data extraction unit 271, a meta parameter learning unit 272, a generalization error evaluation unit 273, and a learning continuation determination unit 274.

[0231] The training data extraction unit 271 extracts the entire acquired data set D optall , extract training data for learning meta-parameter values. The meta parameter learning unit 272 uses the training data extracted by the training data extraction unit 271 to learn the meta parameter values. The generalization error evaluation unit 273 calculates an evaluation value for the generalization error of the predictor when the meta parameter values ​​learned by the meta parameter learning unit 272 are used. The learning continuation determination unit 274 determines whether or not to continue learning of the meta parameter values ​​based on the evaluation value calculated by the generalization error evaluation unit 273.

[0232] FIG. 20 is a diagram showing an example of data input / output in the meta parameter individual processing unit 261 shown in FIG. When the value of the meta-learning execution flag indicates that learning of meta parameter values ​​is to be performed, the training data extraction unit 271 extracts the entire acquired data set D optall The training data extraction unit 271 extracts training data for learning meta parameter values ​​from the meta parameter values. The training data extraction unit 271 repeats the extraction of training data until the value of the meta-learning execution flag reaches a value indicating that continuation of learning is not necessary. The training data extraction unit 271 is an example of a training data extraction means.

[0233] When the value of the meta-learning execution flag indicates that meta-parameter values ​​should be learned, the meta-parameter learning unit 272 learns meta-parameter values ​​based on training data for learning meta-parameter values, learning parameter information, and predictor information. The training data for learning meta-parameter values ​​includes combinations of input values ​​in the learning model and correct answers for the output values ​​of the learning model for those input values. The meta-parameter learning unit 272 is an example of a meta-parameter learning means.

[0234] The predictor information is information about a predictor having a meta-parameter to be learned. For example, the predictor information may include information about a function that represents the predictor. The learning parameter information is information related to meta parameters of the learning target. For example, the learning parameter information may include information indicating the number of meta parameters included in the learning target predictor.

[0235] Here, the predictor to be learned for the meta parameter values ​​is expressed by a function f as shown in equation (26).

[0236]

number

[0237] x denotes the input to the predictor, θ denotes the parameters of the predictor, and y denotes the output of the predictor. The probability distribution p(y|x, θ) of the output of the predictor is expressed as in equation (27).

[0238]

number

[0239] N s is a positive integer indicating the number of parameters of the predictor, and θ = (θ1, θ2, . . . , θ Ns ) Parameter θ i (i=1, 2, . . ., N s ) follows the probability distribution p(θ|S) as shown in equation (28).

[0240]

number

[0241] In the training of a Bayesian neural network, the conditional probability distribution p(θ|S) based on the data S of the parameter θ is calculated. The method by which the learning device 1 obtains the probability distribution p(θ|S) is not limited to a specific method. For example, the learning device 1 may obtain the probability distribution p(θ|S) using the Optimal Gibbs Posterior structure shown in equation (29).

[0242]

number

[0243] P(θ) represents the prior distribution of the value of the parameter θ. The meta parameter learning unit 272 learns this prior distribution P(θ) as the meta parameter value. β is a parameter called a temperature parameter. The value of the temperature parameter β is set in advance, for example.

[0244] "l(S,f(x,θ))" denotes a loss function l based on the difference between the output of the predictor and the correct value of the output based on the correct data S shown as training data. "E" stands for expected value. Specifically, "E θ~P(θ) [exp(-βl(S,f(x,θ)))]" denotes the expected value of "exp(-βl(S,f(x,θ)))" when the parameter θ follows the prior distribution P(θ). The meta parameter learning unit 272 learns the meta parameter values ​​so that the expected value of the loss function shown in equation (30) is as small as possible, for example.

[0245]

number

[0246] "l(S,f θ,P )" represents the loss function l, similar to "l(S,f(x,θ)))" in equation (29). In equation (30), the function f indicating the predictor is defined as "f θ,P " indicates the parameter θ and the probability distribution P, which is a meta-parameter.

[0247] As mentioned above, "E" represents the expected value. S~D [l(S,f θ,P )] represents the expected value of the loss function l when the correct data S follows the probability distribution D. D~Τ [E S~D [l(S,fθ,P )]] is the expected value "E D~Τ [E S~D [l(S,f θ,P )]]" represents the expected value.

[0248] For example, the meta parameter learning unit 272 obtains the probability distribution Q(P) of the probability distribution P(θ) as a meta parameter based on equation (31).

[0249]

number

[0250] “P(P)” represents the prior distribution of the probability distribution P(θ), which is a meta-parameter. λ is a parameter called a temperature parameter, and the value of λ is set in advance, for example. N τ is a positive integer representing the number of tasks. "ln" represents the natural logarithm. As mentioned above, "E" represents the expected value. θ~P(θ) The "[···]" represents the expected value of the value in brackets ([···]) when the value of the parameter θ follows the probability distribution P(θ). "E P~P [···]" represents the expected value of the value in parentheses ([···]) when the probability distribution P(θ) follows the probability distribution P(P).

[0251] The generalization error evaluation unit 273 calculates an evaluation value of the generalization error of the predictor when the above-mentioned probability distributions P(θ) and Q(P) are used. For example, the generalization error evaluation unit 273 calculates an evaluation value of the generalization error L(Q,T) shown in equation (32).

[0252]

number

[0253] As mentioned above, "E" represents the expected value. Specifically, the right-hand side of Equation (32), "E P~Q [E D~Τ[E S~D [l(S,f θ,P )]]]” is the expected value “E” shown in equation (30) when the probability distribution P(θ) follows the probability distribution Q(P). D~Τ [E S~D [l(S,f θ,P )]]" represents the expected value. The generalization error evaluation unit 273 calculates, for example, the value of the right side of equation (33) (the right side of the inequality shown in equation (33)) as the evaluation value of the generalization error L(Q,T).

[0254]

number

[0255] “C(δ,λ,β)” is the loss function l(S,f θ,f ) is a function that is determined depending on the type of The right-hand side of equation (33) represents the upper bound of the generalization error L(Q,T). The right-hand side of equation (33) is also written as L^(Q,T).

[0256] The learning continuation determination unit 274 sets the value of the individual learning continuation flag based on the evaluation value L^(Q,T) of the generalization error calculated by the generalization error evaluation unit 273. The learning continuation determination unit 274 may calculate the value of the individual learning continuation flag I based on equation (34).

[0257]

number

[0258] The value "0" of the individual learning continuation flag I indicates that there is no need to continue learning of meta parameter values. The value "1" of the individual learning continuation flag I indicates that there is a need to continue learning of meta parameter values. ε is a constant indicating a predetermined threshold value.

[0259] The evaluation value L^(Q,T) of the generalization error indicates a smaller value as the evaluation is higher. Therefore, if the value of the evaluation value L^(Q,T) is equal to or less than the threshold ε, the learning continuation determination unit 274 determines that there is no need to continue learning the meta parameter values. On the other hand, if the value of the evaluation value L^(Q,T) is greater than the threshold ε, the learning continuation determination unit 274 determines that it is necessary to continue learning the meta parameter values.

[0260] The learning continuation determination unit 274 may determine whether or not to continue learning of the meta parameter value based on information about the conditions for continuing learning. Fig. 20 shows an example in which the learning continuation determination unit 274 acquires error threshold information and continuation condition information as information about the conditions for continuing learning.

[0261] The error threshold information is a decision threshold for the evaluation value L^(Q,T) of the generalization error, like the threshold ε described above. The continuation condition information is information indicating a determination method other than the determination based on the evaluation value L^(Q,T) of the generalization error. For example, when the number of iterations of learning the meta parameter value reaches a predetermined number, the learning continuation determination unit 274 may determine that it is not necessary to continue learning the meta parameter value even if the evaluation value L^(Q,T) of the generalization error is greater than the threshold ε.

[0262] However, the method by which the learning continuation determination unit 274 determines whether or not it is necessary to continue learning the meta parameter values ​​is not limited to a specific method. The information on the conditions for continuing learning used by the learning continuation determination unit 274 can be various information depending on the method by which the learning continuation determination unit 274 determines whether or not it is necessary to continue learning the meta parameter values.

[0263] 21 is a diagram showing a second example of the configuration of the meta parameter individual processing unit 261. In the example of Fig. 21, the meta parameter individual processing unit 261 includes a meta-learning execution determination unit 281 in addition to the units shown in Fig. 19. The meta-learning execution determination unit 281 sets a meta-learning execution flag.

[0264] FIG. 22 is a diagram showing an example of data input / output in the meta parameter individual processing unit 261 shown in FIG. The meta-learning execution determination unit 281 sets the value of the meta-learning execution flag based on the internal learning evaluation value.

[0265] For example, if the evaluation of the prediction accuracy of the predictor indicated by the internal learning evaluation value is lower than a predetermined evaluation, the meta-learning execution determination unit 281 sets the value of the meta-learning execution flag to a value indicating that learning of the meta-parameter values ​​will be performed. On the other hand, if the evaluation of the prediction accuracy of the predictor indicated by the internal learning evaluation value is higher than or equal to the predetermined evaluation, the meta-learning execution determination unit 281 sets the value of the meta-learning execution flag to a value indicating that learning of the meta-parameter values ​​will not be performed. The meta-learning execution determination unit 281 is an example of a meta-learning execution determination means. In this way, the value of the meta-learning execution flag may be set inside the learning continuation determination unit 274.

[0266] 23 is a diagram showing an example of a skill database update process performed by the learning device 1 according to the third embodiment. For example, the learning device 1 performs the process shown in FIG. 23 when acquiring training data for multiple skills.

[0267] (Step S301) The data update unit 223 updates the entire acquired data set D optall Specifically, the data update unit 223 initializes the entire acquired data set D optall Set the value of to the empty set. After step S301, the process proceeds to step S302.

[0268] (Step S302) The search task setting unit 250 sets a search task. For example, the search task setting unit 250 sets an unknown task parameter value τ j By selecting j The task τ that is associated with j may be set as the search task. After step S302, the process proceeds to step S303.

[0269] Steps S303 to S313 in FIG. 23 are the same as steps S101 to S113 in FIG. The loop from steps S305 to S309 in FIG. 23 is denoted as loop L31.

[0270] In step S313, the high-level controller π H If the high-level controller learning unit 240 determines that learning needs to be continued (step S313: YES), the process proceeds to step S321. On the other hand, the high-level controller π H If the high-level controller learning unit 240 determines that it is not necessary to continue learning (step S313: NO), the process proceeds to step S331.

[0271] Step S321 in FIG. 23 is similar to step S121 in FIG. After step S321, the process returns to step S305. Step S331 in FIG. 23 is the same as step S131 in FIG. After step S331, the process proceeds to step S332.

[0272] (Step S332) The data update unit 223 updates the entire acquired data set D optall As described above, the data update unit 223 updates the generated acquired data set D opt,j The total acquired data set D optall Combine with. After step S332, the process proceeds to step S333.

[0273] (Step S333) The meta parameter processing unit 260 calculates meta parameter values ​​of the predictor. After step S333, the process proceeds to step S334.

[0274] (Step S334) The meta parameter processing unit 260 determines whether or not it is necessary to continue learning the meta parameter values. If the meta parameter processing unit 260 determines that it is necessary to continue learning (step S334: YES), the process proceeds to step S341. On the other hand, if the meta parameter processing unit 260 determines that continuation of learning is not necessary (step S334: NO), the learning device 1 ends the processing of FIG.

[0275] (Step S341) The search task setting unit 250 updates the search task. Specifically, the search task setting unit 250 sets any of the tasks that have not yet been set as a search task as the search task. After step S341, the process returns to step S303.

[0276] 24 is a diagram showing an example of processing for calculating meta parameter values ​​of a predictor by the meta parameter processing unit 260. The meta parameter processing unit 260 performs the processing of FIG. 24 in step S333 of FIG.

[0277] (Step S401) The meta parameter individual processing unit 261 calculates a meta parameter value for each predictor. The meta parameter individual processing unit 261 also determines, for each predictor, whether or not learning of the meta parameter value needs to be continued. The meta parameter individual processing unit 261 may execute the process of step S401 for each predictor in parallel, or may execute the process of step S401 for each predictor sequentially. After the process of step S401 is completed for all the predictors to be processed, the process proceeds to step S402.

[0278] (Step S402) The learning continuation flag integrating unit 262 determines whether or not learning of the meta parameter values ​​for all of the predictors is necessary, based on the determination result of whether or not learning of the meta parameter values ​​for each predictor is necessary. After step S402, the meta parameter processing unit 260 ends the processing of FIG.

[0279] 25 is a diagram showing a first example of processing in which the meta parameter individual processing unit 261 calculates a meta parameter value for each predictor and determines whether or not learning of the meta parameter value needs to be continued. In step S401 of FIG. 24, the meta parameter individual processing unit 261 performs the processing of FIG. 25 for each predictor.

[0280] (Step S411) The training data extraction unit 271 extracts training data for learning meta parameter values ​​from the entire acquisition data set Doptall. After step S411, the process proceeds to step S412.

[0281] (Step S412) The meta parameter learning unit 272 learns the meta parameter values ​​of the predictor being processed. After step S412, the process proceeds to step S413.

[0282] (Step S413) The generalization error evaluation unit 273 calculates an evaluation value of the generalization error when using the meta parameter values ​​obtained by learning. After step S413, the process proceeds to step S414.

[0283] (Step S414) The learning continuation determination unit 274 determines whether or not the learning of the parameter values ​​needs to be continued based on the evaluation value of the generalization error. After step S414, the meta parameter individual processing unit 261 ends the processing of FIG.

[0284] 26 is a diagram showing a second example of processing in which the meta parameter individual processing unit 261 calculates meta parameter values ​​for each predictor and determines whether or not learning of the meta parameter values ​​needs to be continued. In step S401 of FIG. 24, the meta parameter individual processing unit 261 performs the processing of FIG. 26 for each predictor instead of the processing of FIG. 25.

[0285] (Step S421) The meta-learning execution determination unit 281 sets the value of the meta-learning execution flag based on the internal learning evaluation value. After step S421, the process proceeds to step S422.

[0286] Steps S422 to S425 in FIG. 26 are the same as steps S411 to S414 in FIG. After step S425, the meta parameter individual processing unit 261 ends the processing of FIG.

[0287] A more specific example of the skill database update process performed by the learning device 1 according to the third embodiment, shown in FIG. 23, will be described. In step S302, the search task setting unit 250 selects, for example, the shape of an object for which a grasping motion is to be learned as an unknown task parameter. The search task setting unit 250 may sample the unknown task parameter according to a probability distribution T. Alternatively, the search task setting unit 250 may set the unknown task parameter using an algorithm that probabilistically selects the unknown task parameter. The same applies to step S341.

[0288] In step S303, the search point set initialization unit 211 defines a state variable x that represents the position, posture, etc. of the robot 5 and the object to be grasped, and sets the state of the robot 5 and the object to be grasped before the grasping operation is performed as the initial state x si The search point set initialization unit 211 also sets the target state / known task parameter β gi Then, the search point set initialization unit 211 sets the initial state x si and the goal state / known task parameters β gi The pair (x si ,β gi ) into the search point set X ~ search,j Set it to the element.

[0289] In step S306, the system model setting unit 221 sets the search point subset X ~ check Search point X, which is an element of ~ i and extract the target state / known task parameters β gi and the set task τ j Based on this, we will consider the system model (dynamics), constraints in the system model, and the low-level controller π L Examples of the constraint conditions here include, but are not limited to, the operating area of ​​the robot 5, the upper limit of the input in the specifications of the robot 5, and constraint conditions for collision avoidance.

[0290] Furthermore, the system model setting unit 221 sets the search point X ~ i From the initial state x si and the goal state / known task parameters β gi x included in fi Set the following. Furthermore, the system model setting unit 221 sets the evaluation function g in the optimal control problem based on these values. The system model setting unit 221 may set the evaluation function g shown in equation (35).

[0291]

number

[0292] "||·|| 2 " indicates the squared norm. ε g is a tolerance parameter indicating the tolerance of the magnitude of the error.

[0293] In step S312, the prediction accuracy evaluation function setting unit 232 sets the prediction accuracy evaluation function J shown in equation (36) for the predictor configured using the Bayesian neural network. g^i may be set.

[0294]

number

[0295] μ g^j (X ~ ) indicates the predicted mean value. σ g^j 2 (X ~ ) denotes the prediction variance. For Bayesian neural network predictions, these values ​​can be obtained. γ is a coefficient by which the prediction variance is multiplied, and can be understood as a parameter that sets the confidence region (confidence interval). Alternatively, the prediction accuracy evaluation function setting unit 232 may set the level set function g ^ i The function to calculate the entropy of the prediction accuracy evaluation function J g^i It may be set as follows.

[0296] In step S313, the evaluation unit 233 evaluates the search point set X ~ search,j Each element of X ~ Regarding the above prediction variance σ g^j 2 (X ~ ) and calculate σ for all elements. g^j 2 (X ~ )≦ε σ If the following holds, it may be determined that continuation of learning is unnecessary. σ is the prediction variance threshold. σ is also called the dispersion threshold parameter. ~ search,j element of (x si ,β gi ) to X ~ It is written as follows. Or, search point set X ~ search,j For all elements of σ g^j 2 (X ~ )≦ε σ is true, or the acquired data set D opt,j If the number of elements reaches a set threshold, it may be determined that continuation of learning is not necessary.

[0297] As described above, the meta parameter learning unit 272 learns the values ​​of meta parameters that indicate the probability distribution in a learning model in which the parameter values ​​follow a probability distribution, based on training data that indicates the input and output in the learning model. The generalization error evaluation unit 273 calculates an evaluation value indicating an evaluation of the generalization error of the learning model. The learning continuation determination unit 274 determines whether or not learning of the meta parameter values ​​needs to be continued based on the evaluation value indicating the evaluation of the generalization error of the learning model. According to the learning device 1, when learning meta parameter values ​​of a learning model, it is possible to determine whether or not learning needs to be continued, and it is possible to eliminate unnecessary learning, thereby enabling efficient learning.

[0298] Furthermore, the training data extraction unit 271 repeats the selection of training data to be used for learning from the training data for learning the values ​​of meta parameters until it is determined that continuation of learning is not necessary. According to the learning device 1, when learning meta parameter values ​​of a learning model, it is possible to determine whether or not learning needs to be continued, and it is possible to eliminate unnecessary learning, thereby enabling efficient learning.

[0299] Furthermore, the meta-learning execution determination unit 281 determines whether or not to learn the values ​​of the meta parameters based on an evaluation value indicating an evaluation of the generalization error of the learning model. The training data extraction unit 271 selects training data when the meta-learning execution determination unit 281 determines that learning of meta parameter values ​​is to be performed. According to the learning device 1, when learning the meta parameter values ​​of a learning model, it is possible to determine whether or not to continue learning based on an evaluation of the generalization error of the learning model, thereby enabling efficient learning by eliminating unnecessary learning.

[0300] Furthermore, the learning continuation flag integrating unit 262 determines whether or not learning of meta parameter values ​​is required for all of the multiple learning models based on the determination results of each of the multiple learning continuation determination means corresponding to the multiple learning models. According to the learning device 1, it is possible to determine whether or not it is necessary to continue learning meta parameter values ​​for a plurality of learning models, and it is possible to eliminate unnecessary learning, thereby enabling efficient learning.

[0301] In addition, one of the learning models is a high-level controller π that controls the robot 5 to execute modularized tasks. H and the parameter values ​​of the skills are included in the input values ​​for the learning model. The meta parameter learning unit 272 learns the meta parameter values ​​using training data for a plurality of skills. According to the learning device 1, the difference in tasks is dealt with by learning the meta-parameter values, and multiple tasks are handled by a high-level controller π H It can be executed in

[0302] The robot controller 3 also uses a high-level controller π H Equipped with. According to the robot controller 3, the difference in tasks is handled by setting the meta parameter values, and multiple tasks are handled by a high-level controller π using a single learning model. H It can be executed in

[0303] The robot controller 3 also includes a high-level controller π that controls the robot 5 in accordance with the shape of the object to be grasped so that the robot 5 can grasp each of the objects to be grasped having different shapes. H Equipped with. The robot controller 3 is expected to be able to control the robot 5 with high precision according to the shape of the object to be grasped.

[0304] <Fourth embodiment> Fig. 27 is a diagram illustrating an example of the configuration of a learning device according to the fourth embodiment. In the configuration illustrated in Fig. 27, a learning device 610 includes a meta parameter learning unit 611, a generalization error evaluation unit 612, and a learning continuation determination unit 613.

[0305] With this configuration, the meta parameter learning unit 611 learns the values ​​of meta parameters that indicate a probability distribution in a learning model in which the values ​​of the parameters follow a probability distribution, based on training data that indicates the input and output in the learning model. The generalization error evaluation unit 612 calculates an evaluation value indicating an evaluation of the generalization error of the learning model. The learning continuation determination unit 613 determines whether or not learning of the meta parameter values ​​needs to be continued based on the evaluation value indicating the evaluation of the generalization error of the learning model.

[0306] The meta parameter learning unit 611 corresponds to an example of a meta parameter learning means, the generalization error evaluation unit 612 corresponds to an example of a generalization error evaluation means, and the learning continuation determination unit 613 corresponds to an example of a learning continuation determination means. According to the learning device 610, when learning the meta parameter values ​​of a learning model, it is possible to determine whether or not learning needs to continue, and learning can be performed efficiently in that unnecessary learning can be eliminated.

[0307] Fifth Embodiment 28 is a diagram showing an example of the configuration of a control device according to the fifth embodiment. In the configuration shown in FIG. With this configuration, the control unit 621 controls the robot in accordance with the shape of the object to be grasped so that the robot can grasp each of the objects to be grasped having different shapes. The control device 620 is expected to enable highly accurate control of the robot according to the shape of the object to be grasped.

[0308] Sixth Embodiment Fig. 29 is a diagram showing an example of processing in the learning method according to the sixth embodiment. The learning method shown in Fig. 29 includes learning meta parameters (step S611), evaluating generalization errors (step S612), and determining whether to continue learning (step S613).

[0309] In learning meta parameters (step S611), the computer learns the values ​​of meta parameters that indicate the probability distribution in a learning model in which the parameter values ​​follow a probability distribution, based on training data that indicates the input and output in a previous learning model. In evaluating the generalization error (step S612), the computer calculates an evaluation value indicating an evaluation of the generalization error of the learning model.

[0310] In determining whether to continue learning (step S613), the computer determines whether to continue learning the values ​​of the meta parameters based on an evaluation value indicating an evaluation of the generalization error of the learning model. According to the learning method shown in Figure 29, when learning the meta parameter values ​​of a learning model, it is possible to determine whether or not learning needs to continue, and learning can be performed efficiently in that unnecessary learning can be avoided.

[0311] Note that the programs for executing all or part of the processing performed by the learning device 1, robot controller 3, learning device 610, and control device 620 may be recorded on a computer-readable recording medium, and the programs recorded on this recording medium may be read into a computer system and executed to perform the processing of each part. Note that the term "computer system" here includes hardware such as an OS and peripheral devices. Furthermore, "computer-readable recording media" refers to portable media such as flexible disks, optical magnetic disks, ROMs (Read Only Memory), and CD-ROMs (Compact Disc Read Only Memory), as well as storage devices such as hard disks built into computer systems. The program may be one that realizes part of the aforementioned functions, or may be one that can realize the aforementioned functions in combination with a program already stored in the computer system.

[0312] Although an embodiment of the present invention has been described above in detail with reference to the drawings, the specific configuration is not limited to this embodiment, and includes designs within the scope of the gist of the present invention. [Industrial Applicability]

[0313] The present invention may be applied to a learning device, a control device, a learning method, and a recording medium. [Explanation of symbols]

[0314] 1,610 Learning Device 2 Storage device 3 Robot Controller 4. Measuring equipment 5. Robot 100 Control System 210 Search point set setting section 211 Search point set initialization part 212 Order search point set setting part 221 System Model Setting Section 222 Problem setting calculation section 223 Data Update Section 230 Prediction accuracy evaluation function learning unit 231 Level Set Function Learning Unit 232 Prediction accuracy evaluation function setting section 233 Evaluation Department 240 High-level controller learning unit 611 Metaparameter Learning Unit 612 Generalization Error Evaluation Unit 613 Learning Continuation Judgment Unit 620 Control Device 621 Control Unit

Claims

1. a meta parameter learning means for learning values ​​of meta parameters indicating a probability distribution in a learning model in which the values ​​of the parameters follow a probability distribution, based on training data indicating inputs and outputs in the learning model; a generalization error evaluation means for calculating an evaluation value indicating an evaluation of a generalization error of the learning model; a learning continuation determination means for determining whether or not learning of the meta parameter value needs to be continued based on the evaluation value; a learning continuation judgment integration means for judging whether or not learning of the meta parameter values ​​needs to be continued for the plurality of learning models as a whole, based on judgment results of the plurality of learning continuation judgment means corresponding to the plurality of learning models; A learning device comprising:

2. a training data extraction means for repeatedly selecting training data to be used for learning from among the training data for learning the values ​​of the meta parameters until it is determined that continuation of learning is unnecessary; The learning device of claim 1 further comprising:

3. a meta-learning execution determination means for determining whether or not to perform learning of the values ​​of the meta-parameters based on an evaluation value indicating an evaluation of the generalization error of the learning model; Furthermore, the training data extraction means selects the training data when the meta-learning execution determination means determines that learning of the values ​​of the meta parameters is to be performed. The learning device according to claim 2 .

4. one of the learning models is configured as a control means for controlling a control object to execute a task in which the operation of the control object is modularized, and a parameter value of a skill is included in an input value for the learning model; the meta parameter learning means learns the values ​​of the meta parameters using training data of a plurality of skills; The learning device according to any one of claims 1 to 3.

5. A control device comprising the control means according to claim 4.

6. A control means for controlling the robot in accordance with the size of the object to be grasped so that the robot can grasp each of the objects to be grasped having different shapes. The control device of claim 5 , comprising:

7. The computer learning a value of a meta parameter indicating a probability distribution in a learning model in which the value of the parameter follows a probability distribution, based on training data indicating an input and an output in the learning model; calculating an evaluation value indicating an evaluation of a generalization error of the learning model; determining whether or not it is necessary to continue learning the value of the meta parameter based on the evaluation value; determining whether or not to continue learning the values ​​of the meta parameters for all of the learning models based on a determination result of whether or not to continue learning each of the values ​​of the meta parameters corresponding to the learning models; A learning method that includes:

8. On the computer, learning a value of a meta parameter indicating a probability distribution in a learning model in which the value of the parameter follows a probability distribution, based on training data indicating an input and an output in the learning model; calculating an evaluation value indicating an evaluation of a generalization error of the learning model; determining whether or not it is necessary to continue learning the value of the meta parameter based on the evaluation value; determining whether or not to continue learning the meta parameter values ​​for all of the plurality of learning models based on a determination result of whether or not to continue learning each of the plurality of meta parameter values ​​corresponding to the plurality of learning models; A program to execute.

Citation Information

Patent Citations

  • Device and method for driving robot, and robot

    JP2012024882A

  • System and method for controlling actuators of an articulated robot

    JP2020522394A

  • Robust and fast model fitting by adaptive sampling

    US8756175B1

  • System and method for controlling actuators of an articulated robot

    WO2018219943A1

  • Identifying boundaries of lesions within image data

    WO2020201516A1