Control device, control method, control program, learnt model generating device, learnt model generating method and learnt model generating program

The control system uses learned models to estimate posture changes in flexible robots, addressing the challenge of sensor complexity and inaccuracy in contact tasks, thereby enhancing control accuracy and simplifying system configuration.

JP2025185577APending Publication Date: 2025-12-22OMRON CORP
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
JP2024093898
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Filing Date
2024-06-10
Publication Date
2025-12-22

AI Technical Summary

Technical Problem

Controlling the movement of flexible robots during tasks involving contact is challenging due to changes in posture that cannot be accurately detected using existing sensors, leading to complex system configurations and inaccurate control.

Method used

A control system that utilizes supervised and machine learning models to estimate posture changes based on data from simulations, allowing the control of flexible robots without the need for additional sensors.

Benefits of technology

Enables accurate control of flexible robots during contact tasks by estimating posture changes through learned models, simplifying system configuration and improving control accuracy.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2025185577000001_ABST
    Figure 2025185577000001_ABST
Patent Text Reader

Abstract

To control motion of a flexible robot without making a sensor obtain information on a portion whose posture is varied by contacting of flexible robot, when the flexible robot carries out a task accompanied by the contacting.SOLUTION: A control device 14 obtains first information representing an observed state of a real robot, when the real robot that is a flexible robot carries out a task accompanied by contacting thereof. The control device 14 obtains second information by inputting the first information to a first learnt model that outputs the second information that cannot be observed in a real space when the first information is inputted thereto. The control device 14 determines behavior information on the real robot, using a second learnt model for determining behavior of the rear robot and the obtained first information and the obtained second information. The control device 14 controls the real robot so that behavior represented by the determined behavior information can be realized.SELECTED DRAWING: Figure 1
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] The present disclosure relates to a control device, a control method, a control program, a trained model generation device, a trained model generation method, and a trained model generation program. [Background technology]

[0002] Conventionally, there is known a technique for controlling a quadruped robot based on information obtained by simulation (see, for example, Non-Patent Document 1). The technique disclosed in Non-Patent Document 1 controls a quadruped robot using information that can only be obtained by simulation (privileged information).

[0003] Furthermore, when controlling a robot in which one part is more flexible than the other parts (hereinafter simply referred to as a "flexible robot"), a technique is known in which the posture of the flexible robot is detected using motion capture and the flexible robot is controlled (see, for example, non-patent document 2).

[0004] Furthermore, a technique for controlling a flexible robot using the results of a simulation is known (see, for example, Non-Patent Documents 3 and 4). [Prior art documents] [Non-patent literature]

[0005] [Non-Patent Document 1] Lee, J., Hwangbo, J., Wellhausen, L., Koltun, V., & Hutter, M. (2020). Learning quadrupedal locomotion over challenging terrain. Science robotics, 5(47), eabc5986. [Non-patent document 2] Hamaya, M., Lee, R., Tanaka, K., von Drigalski, F., Nakashima, C., Shibata, Y., & Ijiri, Y. (2020, May). Learning robotic assembly tasks with lower dimensional systems by leveraging physical softness and environmental constraints. In 2020 IEEE International Conference on Robotics and Automation (ICRA) (pp. 7747-7753). IEEE. [Non-patent document 3] Jitosho, R., Lum, TGW, Okamura, A., & Liu, K. (2023, December). Reinforcement Learning Enables Real-Time Planning and Control of Agile Maneuvers for Soft Robot Arms. In Conference on Robot Learning (pp. 1131-1153). PMLR. [Non-patent document 4] Graule, MA, McCarthy, TP, Teeple, CB, Werfel, J., & Wood, RJ (2022). SoMogym: A toolkit for developing and evaluating controllers and reinforcement learning algorithms for soft robots. IEEE Robotics and Automation Letters, 7(2), 4071-4078. Summary of the Invention [Problem to be solved by the invention]

[0006] When a flexible robot performs a task involving contact, the posture of the flexible robot changes when the flexible robot itself or an object held by the flexible robot comes into contact with another object. In such cases, in order to accurately control the movement of the flexible robot, it is necessary to sequentially detect changes in the posture of the flexible robot due to contact.

[0007] When controlling a flexible robot, sensors for sequentially detecting positions where the robot's posture changes due to contact with other objects can be, for example, motion capture devices or cameras. However, incorporating such sensors into a robot system complicates the system configuration. Furthermore, even if the flexible robot is controlled using data obtained from such sensors, it is difficult to accurately control the movement of the flexible robot if the accuracy of the data obtained by the sensors is low or if the posture change of the flexible robot is small.

[0008] The present disclosure has been made in consideration of the above points, and aims to control the operation of a flexible robot when it performs a task involving contact, without using sensors to obtain information about areas where the posture changes due to contact. [Means for solving the problem]

[0009] In order to achieve the above-mentioned object, the control device of the present disclosure includes: a first acquisition unit that acquires first information representing an observed state of a real robot, which is a robot in real space and has some parts that are more flexible than other parts, when the real robot performs a task involving contact; a second acquisition unit that acquires second information by inputting the first information acquired by the first acquisition unit into a first learned model that outputs second information that is not observed in the real space when the first information is input; a second learned model for determining the behavior of the real robot; a determination unit that determines behavior information of the real robot using the first information acquired by the first acquisition unit and the second information acquired by the second acquisition unit; and a control unit that controls the real robot so that the behavior represented by the behavior information determined by the determination unit is realized, wherein the first learned model is a model obtained by supervised machine learning based on data obtained by running a simulation in which a virtual robot, which is a robot on a computer corresponding to the real robot, performs the task, and the second learned model is a model obtained by machine learning based on data obtained by running the simulation.

[0010] The control method disclosed herein is a control method in which a computer executes processing, in which the computer acquires first information representing an observed state of a real robot, the real robot being a robot in real space where some parts of the robot are more flexible than other parts, when the real robot performs a task involving contact, acquires second information by inputting the acquired first information into a first learned model that outputs second information that is not observed in the real space when the first information is input, acquires second information by inputting the acquired first information into a second learned model for determining the behavior of the real robot, determines behavior information of the real robot using the acquired first information and the acquired second information, and controls the real robot so that the behavior represented by the determined behavior information is realized, wherein the first learned model is a model obtained by supervised machine learning based on data obtained by running a simulation in which a virtual robot, which is a robot on a computer corresponding to the real robot, performs the task, and the second learned model is a model obtained by machine learning based on data obtained by running the simulation.

[0011] The control program disclosed herein is also a control program for causing a computer to execute processing, which includes acquiring first information representing an observed state of a real robot, which is a robot in real space and has some parts that are more flexible than other parts, when the real robot performs a task involving contact, acquiring the second information by inputting the first information acquired by the acquisition unit into a first learned model that outputs second information that is not observed in the real space when the first information is input, determining behavioral information of the real robot using the acquired first information and the acquired second information, and controlling the real robot so that the behavior represented by the determined behavioral information is realized, wherein the first learned model is a model obtained by supervised machine learning based on data obtained by running a simulation in which a virtual robot, which is a robot on a computer corresponding to the real robot, performs the task, and the second learned model is a model obtained by machine learning based on data obtained by running the simulation.

[0012] The trained model generation device disclosed herein is a trained model generation device that includes a simulation unit that executes a simulation of a virtual robot on a computer, where some parts of the robot are more flexible than other parts, performing a task involving contact; a training data acquisition unit that acquires, from data obtained while the simulation is being executed, first training information that represents the observed state of the virtual robot, second training information that is not observed in a real space in which a real robot corresponding to the virtual robot exists, and training behavior information that represents the behavior of the virtual robot; and a learning unit that uses supervised machine learning to generate a first trained model that outputs the second information when the first information is input, based on the first training information, second training information, and training behavior information acquired by the training data acquisition unit, and uses machine learning to generate a second trained model for determining the behavior of the real robot.

[0013] The trained model generation method disclosed herein is a trained model generation method in which a computer executes a simulation in which a virtual robot on a computer, where some parts of the robot are more flexible than other parts, performs a task involving contact, and from data obtained while the simulation is being executed, first learning information representing the observed state of the virtual robot, second learning information that is not observed in a real space in which a real robot corresponding to the virtual robot exists, and learning behavior information representing the behavior of the virtual robot are obtained, and based on the obtained first learning information, second learning information, and learning behavior information, supervised machine learning is used to generate a first trained model that outputs the second information when the first information is input, and machine learning is used to generate a second trained model for determining the behavior of the real robot.

[0014] The trained model generation program disclosed herein is a trained model generation program for causing a computer to execute a process of: executing a simulation of a virtual robot on a computer, where some parts of the robot are more flexible than other parts, performing a task involving contact; acquiring, from data obtained while the simulation is being executed, first learning information representing the observed state of the virtual robot, second learning information that is not observed in a real space in which a real robot corresponding to the virtual robot exists, and learning behavior information representing the behavior of the virtual robot; and, based on the acquired first learning information, second learning information, and learning behavior information, using supervised machine learning to generate a first trained model that outputs the second information when the first information is input, and using machine learning to generate a second trained model for determining the behavior of the real robot. [Effects of the Invention]

[0015] According to the control device, control method, control program, trained model generation device, trained model generation method, and trained model generation program disclosed herein, when a flexible robot performs a task involving contact, the operation of the flexible robot can be controlled without using a sensor to obtain information about the areas where the posture changes due to contact. [Brief explanation of the drawings]

[0016] [Figure 1] FIG. 1 is a diagram for explaining the present embodiment. [Figure 2] FIG. 2 is a diagram for explaining a control system according to the present embodiment. [Figure 3] FIG. 1 is a diagram for explaining a trained model of the present embodiment. [Figure 4] 1 is a block diagram showing a schematic configuration of a control system according to an embodiment of the present invention; [Figure 5] FIG. 2 is a block diagram showing the hardware configuration of the control device according to the present embodiment. [Figure 6] 1 is a flowchart showing the flow of a trained model generation process in this embodiment. [Figure 7] 4 is a flowchart showing the flow of a control process in the present embodiment. [Figure 8] FIG. 10 is a diagram showing the results of this example. [Figure 9] FIG. 10 is a diagram for explaining a modified example of the present embodiment. DETAILED DESCRIPTION OF THE INVENTION

[0017] An example of an embodiment of the present disclosure will be described below with reference to the drawings. In this embodiment, a control system equipped with a control device according to the present disclosure will be described as an example. Note that the same reference numerals are used in the drawings to designate identical or equivalent components and parts. Furthermore, the dimensions and proportions of the drawings are exaggerated for the sake of explanation and may differ from the actual proportions.

[0018] Fig. 1 is a diagram for explaining this embodiment. As shown in Fig. 1, a flexible robot 12S in a virtual space Sim (hereinafter simply referred to as a virtual robot) and a flexible robot 12R in a real space Real (hereinafter simply referred to as a real robot) perform a peg-in-hole task such as inserting pegs P_S and P_R into holes.

[0019] As shown in Fig. 1, the real robot 12R includes an arm 12R_A, a gripper 12R_G, and a flexible portion 12R_S. As shown in Fig. 1, a flexible portion 12R_S is provided between the arm 12R_A and the gripper 12R_G of the real robot 12R, and the presence of this flexible portion 12R_S enables the real robot 12R to smoothly perform tasks that involve contact. Like the real robot 12R, the virtual robot also includes an arm 12S_A, a gripper 12S_G, and a flexible portion 12S_S.

[0020] As shown in FIG. 1, in this embodiment, a simulation is performed in which a virtual robot 12S in a virtual space Sim performs a task involving contact, and a trained model is generated for controlling the behavior of a real robot 12R when performing the same task.

[0021] In this embodiment, it is possible to observe the position information of the arm 12R_A of the real robot 12R and the force information acting on the arm 12R_A. On the other hand, since the flexible portion 12R_S is present between the arm 12R_A and the gripper 12R_G, it is not possible to observe the position information of the peg P_R gripped by the gripper 12R_G.

[0022] Therefore, in this embodiment, a simulation is performed in which the virtual robot 12S performs the same task as the real robot 12R. Then, using privileged information obtained during the simulation that is not observed in the real space, a trained model is generated for estimating the privileged information from information obtained in the real space. This trained model corresponds to a student encoder model, which will be described later. The privileged information is information that cannot be observed unless a sensor, such as a motion capture device or a camera, is added to the system.

[0023] By controlling the real robot 12R in real space using a trained model for estimating hidden information, it is possible to control the behavior of the real robot 12R without using a special sensor for observing the hidden information.

[0024] A specific description will be given below. In the following, when there is no distinction between real space and virtual space, the components will be simply referred to as the "robot 12," the "arm 12A," the "flexible portion 12S," the "gripping portion 12G," and the "peg P." The task performed by the robot 12 of this embodiment is a task involving contact. More specifically, the task of this embodiment is a task in which the robot 12 itself or the peg P, which is an example of an implement grasped by the robot 12, needs to operate while coming into contact with another object. For example, a peg-in-hole task is a task in which the peg P needs to operate while coming into contact with the floor or base on which the hole H is located.

[0025] <Framework of this embodiment> (A. Problem setting) 2 is a diagram for explaining an outline of the control system of this embodiment. As shown in FIG. 2, in this embodiment, a peg-in-hole task is executed in which the robot 12 inserts a peg P into a hole H. As shown in FIG. 2, in this embodiment, posture information p wrist and force information F acting on arm 12A wrist and a vector containing observation information ot 2, in this embodiment, the relative position information p peg and the attitude angle information p of peg P θ and alignment signal b align and a vector containing the secret information x t Let's say.

[0026] (B. Overview) 3 is a diagram for explaining the trained model of this embodiment. As shown in FIG. 3, in this embodiment, the secret information x t and observation information t When the input is made, the behavior information a representing the behavior of the robot 12 is generated. t The teacher policy model M2 outputs the observed information o t When input, the estimated value of the secret information x - t A student encoder model M1 that outputs

[0027] Specifically, as shown in Fig. 3, a trained teacher policy model M2 is generated by performing reinforcement learning based on data obtained by simulating the virtual robot 12S performing a task. When the simulation is performed, domain randomization is also performed to vary the simulation parameters (e.g., the size of the hole or the softness of the flexible part).

[0028] 3, a student encoder model M1 is generated by performing a known supervised machine learning based on data obtained by performing a simulation in which the virtual robot 12S performs a task. In this case, as shown in FIG. 3, the secret information x t and the secret information estimate x output from the student encoder model M1. - t Known supervised machine learning is performed to reduce the loss representing the difference between

[0029] This student encoder model M1 uses the observation information o t When input, the estimated value of the secret information x - t Therefore, by using the student encoder model M1, it is possible to obtain the estimated value x of the secret information that cannot be observed in the real space. - t When controlling the movement of the real robot 12R, the observation information detected in the real space is used. t By inputting this to the trained student encoder model M1, the secret information estimate x - t and estimate the secret information x - t and observation information detected in real space. t and are input to the trained teacher policy model M2 to obtain the action information a t Then, behavioral information a t The behavior of the real robot 12R is controlled according to the received data.

[0030] (C. Learning teacher strategy models) Next, the reinforcement learning method of the above-mentioned teacher policy model M2 will be explained in detail.

[0031] (1. Status) In this embodiment, the state in reinforcement learning is the observed information o t Observation information o t The posture information p of the arm 12A contained in wrist contains three-dimensional posture information, and the force information F acting on the arm 12A wrist contains three-dimensional force information. Therefore, the observation information o t is a six-dimensional vector. The posture information p wrist may be calculated based on values ​​detected by encoders installed at the respective joints of the real robot 12R. wrist The posture information p may be acquired from a force sensor installed on the arm 12R_A.wrist and force information F wrist may further include rotation information. In this case, the orientation information p wrist and force information F wrist becomes six-dimensional information.

[0032] On the other hand, confidential information x t is a 10-dimensional vector. t The relative position information p of peg P contained in peg is information that represents the difference between the position of the peg P in the x, y, and z directions and the position of the hole H. Therefore, the relative position information p peg is three-dimensional information. Also, the secret information x t The attitude angle information p of peg P contained in θ is six-dimensional information expressed by a rotation matrix. t Alignment signal b contained in align takes a value between -1 and 1. As will be described later, the alignment signal b align is a control value that becomes 1 when the error between the position of the peg P and the position of the hole H is less than a predetermined value.

[0033] (2. Action) In this embodiment, the behavior in reinforcement learning is behavior information a t Equivalent to: Behavioral information a t is expressed by the following formula:

[0034]

number

[0035] Therefore, behavioral information a t is the posture information p of the arm 12A. wrist This is information that represents changes in

[0036] (3. Rewards) In this embodiment, the following reward function r is used as the reward function in reinforcement learning.

[0037]

number

[0038] r in the above formula (1) p is the progress reward, which is the reward related to the difference between the position of the tip of the peg P and the position of the hole H. p is expressed by the following formula:

[0039]

number

[0040] The above e t is a vector that represents the difference between the position of the tip of the peg P and the position of the hole H in the x, y, and z directions. W above represents a weighting matrix, and for example, if importance is placed on the difference in position in the z-axis direction, W=[1,1,10] is set.

[0041] r in the above formula (1) i is the reward for inserting peg P, and is expressed as follows:

[0042]

number

[0043] In the above formula (a t ) z represents the action in the z-axis direction, and w i is a function that is 0.001 when the alignment signal described below is 1, and is 1 in other situations.

[0044] r in the above formula (1) a is the reward for smoothness of behavior, and is expressed by the following formula:

[0045]

number

[0046] r in the above formula (1) s is the reward for the task success, which is a function that is 1 when the peg P is inserted into the hole H and is 0 in other situations. For example, the above ||e t If ||<0.005, the peg P is considered to be inserted into the hole H.

[0047] (4. Alignment Signal) Alignment signal b as described above align is a function expressed by the following equation:

[0048]

number

[0049] By training the teacher policy model M2 using the above reward function r, the behavioral information a for smoothly inserting the peg P into the hole H is obtained. t It is possible to obtain a model that outputs

[0050] (C. Learning the Student Encoder Model) Next, a learning method for the student encoder model M1 will be described.

[0051] In this embodiment, the student encoder model M1 is generated by known supervised machine learning based on the simulation data used to train the teacher policy model M2. Specifically, when the terminal time T is set to 20, the observed information o obtained in 1 second is t Time series data [o t-20 ,···o t-1 ] is input to the student encoder model M1, and the student encoder model M1 estimates the secret information x - t will be output.

[0052] When the student encoder model M1 is trained using supervised machine learning, the student encoder model M1 is trained so that the following loss function L becomes small:

[0053]

number

[0054] In addition, the above x p is confidential information x t The relative position information p of peg P peg and the attitude angle information p of peg P θ Therefore, L mse is the correct information obtained in the simulation x p and the estimated value x output from the student encoder model M1. - p is the mean square error between L bce is a binary cross-entropy loss function, and the ground truth alignment signal b align and the estimated alignment signal b output from the student encoder model M1. align - The w in the above is a weighting coefficient, and is set to, for example, w=0.1.

[0055] <Control System 10> Fig. 4 is a block diagram showing a schematic configuration of a control system 10 according to this embodiment. As shown in Fig. 4, the control system 10 includes a sensor group 11, a real robot 12R, and a control device 14. The control device 14 generates a teacher policy model M2 and a student encoder model M1 for controlling the movement of the gripper 12R_G of the real robot 12R. The control device 14 also controls the movement of the real robot 12R using the generated trained teacher policy model M2 and trained student encoder model M1.

[0056] The sensor group 11 is attached to a predetermined location of the real robot 12R, and receives the above-mentioned observation information o tThe observation information o t is an example of the first information of the present disclosure. t to the control device 14. For example, the sensor group 11 includes the above-mentioned force sensor and encoder.

[0057] Fig. 5 is a block diagram showing the hardware configuration of the control device 14 according to this embodiment. As shown in Fig. 5, the control device 14 has a CPU (Central Processing Unit) 42, a memory 44, a storage device 46, an input / output I / F (Interface) 48, a storage medium reader 50, and a communication I / F 52. Each component is connected to each other via a bus 54 so as to be able to communicate with each other.

[0058] The storage device 46 stores a trained model generation program and a control program for executing each process described below. The CPU 42 is a central processing unit that executes various programs and controls each component. That is, the CPU 42 reads the program from the storage device 46 and executes the program using the memory 44 as a work area. The CPU 42 controls each component and performs various arithmetic operations in accordance with the program stored in the storage device 46.

[0059] The memory 44 is made up of RAM (Random Access Memory) and serves as a working area to temporarily store programs and data. The storage device 46 is made up of ROM (Read Only Memory), HDD (Hard Disk Drive), SSD (Solid State Drive), etc., and stores various programs including the operating system and various data.

[0060] The input / output I / F 48 is an interface for inputting data from an external device and outputting data to an external device. Also, input devices for various inputs, such as a keyboard or a mouse, and output devices for various information outputs, such as a display or a printer, may be connected. A touch panel display may be used as the output device, allowing it to function as an input device.

[0061] The storage medium reader 50 reads data stored in various storage media such as CD (Compact Disc)-ROM, DVD (Digital Versatile Disc)-ROM, Blu-ray Disc, and USB (Universal Serial Bus) memory, and writes data to the storage media.

[0062] The communication I / F 52 is an interface for communicating with other devices, and uses standards such as Ethernet (registered trademark), FDDI, and Wi-Fi (registered trademark).

[0063] Next, the functional configuration of the control device 14 will be described. As shown in Fig. 4, the control device 14 functionally includes a simulation unit 18, a learning data acquisition unit 20, a learning unit 22, a first acquisition unit 24, a second acquisition unit 26, a determination unit 28, and a control unit 30. A predetermined storage area of ​​the control device 14 also includes a data storage unit 32 and a trained model storage unit 34. Each functional configuration is realized by the CPU 42 reading out each program stored in the storage device 46, expanding it into the memory 44, and executing it.

[0064] The data storage unit 32 stores the observation information detected by the sensor group 11. t The data storage unit 32 also stores control data used to control the real robot 12R.

[0065] The trained model storage unit 34 stores a trained teacher policy model M2 and a trained student encoder model generated by the process described below.

[0066] First, the simulation unit 18, the learning data acquisition unit 20, and the learning unit 22 generate a trained teacher policy model M2 and a trained student encoder model for controlling the movement of the real robot 12R.

[0067] First, the simulation unit 18 executes a simulation in which a virtual robot 12S on a computer performs a task involving contact. The virtual robot 12S corresponds to a real robot 12R existing in real space, and is a flexible robot having one part more flexible than the other parts, just like the real robot 12R.

[0068] The learning data acquisition unit 20 acquires learning observation information o representing the observation state of the virtual robot 12S from the data obtained during the execution of the simulation. t and secret information x for learning that is not observed in the real space where the real robot 12R corresponding to the virtual robot 12S exists. t and learning behavior information a that represents the behavior of virtual robot 12S. t and get.

[0069] The learning unit 22 uses the learning observation information o acquired by the learning data acquisition unit 20. t and the secret information for training x t and learning behavior information a t Based on the above, the learning unit 22 generates a teacher policy model M2 for determining the behavior of the real robot 12R by using reinforcement learning, which is an example of machine learning. t and the secret information for training x t and learning behavior information a t Based on this, we use supervised machine learning to t When is entered, secret information x t Generate a student encoder model M1 that outputs

[0070] Specifically, the learning unit 22 generates a student encoder model M1 and a teacher strategy model M2 using the learning method described above.

[0071] The task in this embodiment is a peg-in-hole task in which the real robot 12R or the virtual robot 12S inserts a peg held by the real robot 12R or the virtual robot 12S into a hole. The teacher policy model M2 is a model generated by reinforcement learning based on a reward function r such that the reward increases as the distance between the peg held by the virtual robot 12S and the hole decreases. Therefore, the teacher policy model M2 in this embodiment is a model generated by reinforcement learning based on a reward function such that the reward increases as the distance between the current position of the robot 12 itself or a peg, which is an example of a tool held by the robot 12, and the hole, which is an example of a destination of a task that requires movement while coming into contact with other objects, decreases. The determination unit 28 and the control unit 30, which will be described later, use this teacher policy model M2 to control the movement of the real robot 12R.

[0072] The trained student encoder model M1 is an example of a first trained model of the present disclosure. The trained teacher policy model M2 is an example of a second trained model of the present disclosure. The secret information x t is an example of the second information of the present disclosure. Therefore, the student encoder model M1 is a model obtained by supervised machine learning based on data obtained by executing a simulation in which a virtual robot 12S, which is a robot on a computer corresponding to the real robot 12R, performs a task. Also, the teacher policy model M2 is a model obtained by machine learning based on data obtained by executing a simulation. Also, the teacher policy model M2 is a model obtained by reinforcement learning based on data obtained by executing a simulation, and is obtained by using observation information o t and confidential information x t The behavior information a of the robot 12R is input. tThe student encoder model M1 and the teacher strategy model M2 are realized by known machine learning models. For example, the student encoder model M1 is realized by a temporal convolutional network. Also, for example, the teacher strategy model M2 is realized by a multi-layer perceptron.

[0073] Then, the learning unit 22 stores the learned student encoder model M1 and the learned teacher policy model M2 in the learned model storage unit .

[0074] Once the trained student encoder model M1 and the trained teacher policy model M2 are stored in the trained model storage unit 34, it becomes possible to use these models to control the behavior of the real robot 12R. Therefore, the first acquisition unit 24, the second acquisition unit 26, and the control unit 30 use the trained student encoder model M1 and the trained teacher policy model M2 stored in the trained model storage unit 34 to control the behavior of the real robot 12R.

[0075] The first acquisition unit 24 acquires observation information o representing the observation state of the real robot 12R when the real robot 12R performs a task involving contact. t Get.

[0076] The second acquisition unit 26 reads out the trained student encoder model M1 from the trained model storage unit 34. Then, the second acquisition unit 26 performs the observation information o acquired by the first acquisition unit 24 on the trained student encoder model M1. t By entering the secret information x t Get.

[0077] The determination unit 28 reads out the teacher policy model M2 for determining the behavior of the real robot 12R from the learned model storage unit 34. Then, the determination unit 28 reads out the teacher policy model M2 for determining the behavior of the real robot 12R from the learned model storage unit 34. t and the secret information x acquired by the second acquisition unit 26. t Using this, the behavior information of the real robot 12R is t Determine.

[0078] Specifically, the determination unit 28 applies the observation information o acquired by the first acquisition unit 24 to the trained teacher policy model M2. t and the secret information x acquired by the second acquisition unit 26. t By inputting the above, the action information a output from the trained teacher policy model M2 is t Get.

[0079] The control unit 30 determines the behavior information a determined by the determination unit 28. t The real robot 12R is controlled so that the behavior represented by is realized.

[0080] Next, the operation of the control system 10 according to this embodiment will be described.

[0081] When the control device 14 receives a predetermined instruction signal, the CPU 42 of the control device 14 reads out the trained model generation program from the storage device 46, loads it into the memory 44, and executes it. As a result, the CPU 42 functions as each functional component of the control device 14, and the trained model generation process shown in FIG. 6 is executed.

[0082] In step S100, the simulation unit 18 executes a simulation in which the virtual robot 12S on the computer performs a task involving contact.

[0083] In step S102, the learning data acquisition unit 20 acquires learning data from the data obtained during the execution of the simulation in step S100. The learning data in this embodiment is learning observation information o representing the observation state of the virtual robot 12S. t and secret information x for learning that is not observed in the real space where the real robot 12R corresponding to the virtual robot 12S exists. t and learning behavior information a that represents the behavior of virtual robot 12S. t This is data representing a combination of

[0084] In step S104, the learning unit 22 generates a teacher policy model M2 for determining the behavior of the real robot 12R by using reinforcement learning, which is an example of machine learning, based on the learning data acquired in step S102. Also, the learning unit 22 generates a teacher policy model M2 for determining the behavior of the real robot 12R by using supervised machine learning based on the learning data acquired in step S102. t When is entered, secret information x t Generate a student encoder model M1 that outputs

[0085] In step S106, the learning unit 22 stores the trained student encoder model M1 and the trained teacher policy model M2 in the trained model storage unit .

[0086] Next, when the control device 14 receives a predetermined instruction signal, the control device 14 executes the control process shown in Fig. 7. The control process in Fig. 7 is executed repeatedly.

[0087] In step S200, the first acquisition unit 24 acquires observation information o representing the observation state of the real robot 12R when the real robot 12R performs a task involving contact. t Get.

[0088] In step S202, the second acquisition unit 26 reads out the trained student encoder model M1 from the trained model storage unit 34. Then, in step S202, the second acquisition unit 26 applies the observation information o acquired in step S200 to the trained student encoder model M1. t By entering the secret information x t Get.

[0089] In step S204, the determination unit 28 reads out the trained teacher policy model M2 from the trained model storage unit 34. Then, in step S204, the determination unit 28 applies the observation information o acquired in step S200 to the trained teacher policy model M2. t and the secret information x obtained in step S204. tBy inputting the above, the action information a output from the trained teacher policy model M2 is t Get.

[0090] In step S206, the control unit 30 calculates the behavior information a acquired in step S204. t The real robot 12R is controlled so that the behavior represented by is realized.

[0091] As described above, the control device according to this embodiment acquires observation information representing the observation state of the real robot 12R when the real robot 12R performs a task involving contact. The control device acquires the hidden information by inputting the acquired observation information to a trained student encoder model that, when input with the observation information, outputs hidden information that is not observed in real space. The control device determines behavior information for the real robot 12R using a trained teacher policy model for determining the behavior of the real robot 12R, the acquired observation information, and the acquired hidden information. The control device controls the real robot 12R so as to realize the behavior represented by the determined behavior information. Note that the student encoder model according to this embodiment is a model obtained by supervised machine learning based on data obtained by executing a simulation in which a virtual robot 12S, a computer-based robot corresponding to the real robot 12R, performs a previous task. Furthermore, the teacher policy model according to this embodiment is a model obtained by machine learning based on data obtained by executing a simulation. As a result, when the flexible robot performs a task involving contact, the behavior of the flexible robot can be controlled without acquiring, using sensors, information on locations whose posture changes due to contact.

[0092] Furthermore, a control device according to this embodiment executes a simulation in which a virtual robot on a computer performs a task involving contact. The control device acquires, from data obtained during the simulation, learning observation information representing the observation state of the virtual robot, learning secret information not observed in a real space in which a real robot corresponding to the virtual robot exists, and learning behavior information representing the behavior of the virtual robot. Based on the acquired learning observation information, learning secret information, and learning behavior information, the control device uses supervised machine learning to generate a student encoder model that outputs secret information when observation information is input, and also uses machine learning to generate a teacher policy model for determining the behavior of the real robot. This makes it possible to obtain a trained model for controlling the movement of a flexible robot when the flexible robot performs a task involving contact, without using sensors to acquire information on locations whose posture changes due to contact.

[0093] Furthermore, the control device of this embodiment makes it possible to obtain a trained model for enabling a flexible robot to perform a task without collecting real-world data. Furthermore, since it is not necessary to measure the posture of the flexible robot's hand (the part connected by the flexible portion), sensors such as motion capture or cameras are not required. [Example]

[0094] Next, an example will be described. In this example, an experiment was conducted to verify the effectiveness of the proposed method. In this experiment, an experiment was conducted on a peg-in-hole task.

[0095] Figure 8 shows the experimental results. The horizontal axis of Figure 8 represents the shape of the hole in the peg-in-hole task, and the vertical axis represents the task success rate. "Student" in Figure 8 represents the results corresponding to the proposed method described above. "Student (no alignment)" in Figure 8 represents the results when the alignment signal was not used in the proposed method described above, and "TCN" in Figure 8 represents the results when the method disclosed in Reference 1 below was used.

[0096] Reference 1: J. Lee, J. Hwangbo, L. Wellhausen, V. Koltun, and M. Hutter, “Learning quadrupedal locomotion over challenging terrain,” Science robotics, vol. 5, no. 47, p. eabc5986, 2020.

[0097] As can be seen from Figure 8, by using the method of this embodiment, when a flexible robot performs a task involving contact, it is possible to control the movement of the flexible robot without using sensors to obtain information about the locations where the posture changes due to contact.

[0098] The technology of the present disclosure is not limited to the above-described embodiment, and various modifications and applications are possible within the scope of the gist of this disclosure.

[0099] For example, in the above embodiment, the case where the task involving contact by the flexible robot is a peg-in-hole task has been described as an example, but the present invention is not limited to this. The above embodiment can be applied to any task that involves contact by the flexible robot. For example, the above embodiment can be applied to a task in which a flexible robot holds a pen and writes using the pen, a task in which a flexible robot holds a file and polishes using the file, a task in which a flexible robot wipes dirt off a table using a cloth or cleaner, or a task in which a flexible robot cuts ingredients using a knife.

[0100] In the above embodiment, reinforcement learning and supervised learning are used as machine learning algorithms. However, these algorithms may be any algorithms. For example, examples of reinforcement learning algorithms include proximal policy optimization (PPO), soft-actor-critic (SAC), deep deterministic policy gradient (DDPG), and twin-delayed DDPG (TD3). Any algorithm may also be used as the supervised learning algorithm. In the above embodiment, the student encoder model M1 is implemented by a temporal convolutional network, and the teacher policy model M2 is implemented by a multi-layer perceptron. However, this is not a limitation. The student encoder model M1 and the teacher policy model M2 may be implemented by any models. For example, the student encoder model M1 may be implemented by a long short-term memory (LSTM), a recurrent network, or a transformer.

[0101] In the above embodiment, the reward function of the above formula (1) is used as the reward function for reinforcement learning, but the present invention is not limited to this and any reward function may be used. For example, a sparse reward function may be used that is 1 only when the peg P enters the hole H and is 0 otherwise.

[0102] In the above embodiment, a case where a teacher policy model is generated using a reinforcement learning algorithm as a machine learning algorithm has been described as an example, but the present invention is not limited to this. Other machine learning algorithms (e.g., self-supervised learning algorithms) may be used to generate a learned model. Fig. 9 shows a modified example of the teacher policy model. The teacher policy model M2 shown in Fig. 9 is a model that has undergone self-supervised learning based on data obtained by executing a simulation. The teacher policy model M2 in Fig. 9 is generated based on observation information o at time t. t and secret information x at time tt and behavioral information a at time t t When the candidate of is input, the observation information o at time t+1 t and the secret information x at time t+1 t In this case, the determination unit 28 outputs the observation information o at time t acquired by the first acquisition unit 24 to the teacher policy model M2 in FIG. t and the secret information x at time t acquired by the second acquisition unit 26 t and behavioral information a at time t t The secret information x at time t+1 output from the teacher policy model M2 is input. t Based on this, the behavioral information a at time t t Among the candidates, the behavior information of the real robot 12R is t In this case, the secret information x t contains information representing the positional relationship between the peg P and the hole H. Therefore, the determining unit 28 uses the behavior information a t Among the candidates, the behavior information of the real robot 12R is t When determining the secret information x t Based on the information representing the positional relationship included in t Specifically, the behavior information a is determined so that the positional relationship in which the peg P is closest to the hole H is realized. t is decided as a formal action.

[0103] In the above embodiment, the control device 14 executes both the trained model generation process of Fig. 6 and the control process of Fig. 7, but the present invention is not limited to this. For example, a trained model generation device implemented by a computer separate from the control device 14 may be provided, and the trained model generation device may execute the trained model generation process of Fig. 6, while the control device 14 executes the control process of Fig. 7. In this case, the trained model generation device includes at least the simulation unit 18, the learning data acquisition unit 20, and the learning unit 22 described above.

[0104] Furthermore, the processes executed by the CPU after reading the software (program) in the above-described embodiments may be executed by various processors other than the CPU. Examples of such processors include programmable logic devices (PLDs) whose circuit configuration can be changed after fabrication, such as field-programmable gate arrays (FPGAs), and dedicated electrical circuits, such as application-specific integrated circuits (ASICs), which are processors with circuit configurations specifically designed to execute specific processes. Each process may be executed by one of these various processors, or by a combination of two or more processors of the same or different types (e.g., multiple FPGAs, or a combination of a CPU and an FPGA). The hardware structure of these various processors is, more specifically, an electrical circuit that combines circuit elements such as semiconductor devices.

[0105] In the above embodiment, the programs are pre-stored (installed) in a storage device, but the present invention is not limited to this. The programs may be provided in a form stored on a storage medium such as a CD-ROM, DVD-ROM, Blu-ray disc, or USB memory. The programs may also be downloaded from an external device via a network.

[0106] (Addendum) The following additional notes are provided regarding aspects of the present disclosure.

[0107] (Appendix 1) a first acquisition unit that acquires first information representing an observation state of a real robot, the real robot being a robot in real space and having a part of the robot more flexible than other parts, when the real robot performs a task involving contact; a second acquisition unit that acquires second information by inputting the first information acquired by the first acquisition unit into a first trained model that outputs second information that is not observed in the real space when the first information is input; a determination unit that determines behavior information of the real robot using a second learned model for determining the behavior of the real robot, the first information acquired by the first acquisition unit, and the second information acquired by the second acquisition unit; a control unit that controls the real robot so that the behavior represented by the behavior information determined by the determination unit is realized, the first trained model is a model obtained by supervised machine learning based on data obtained by executing a simulation in which a virtual robot, which is a robot on a computer corresponding to the real robot, performs the task; and The second trained model is a model that has been machine-learned based on data obtained by executing the simulation. Control device. (Appendix 2) The task is a task that requires the robot itself or an instrument held by the robot to perform an operation while coming into contact with another object. 10. The control device of claim 1. (Appendix 3) the second trained model is a model that has undergone reinforcement learning based on data obtained by executing the simulation, and that outputs behavior information of the real robot when the first information and the second information are input; the determination unit acquires the behavioral information by inputting the first information acquired by the first acquisition unit and the second information acquired by the second acquisition unit to the second trained model. 10. The control device according to claim 1 or 2. (Appendix 4) the task is a peg-in-hole task in which the real robot or the virtual robot inserts a peg held by the real robot or the virtual robot into a hole; The second trained model is a model that has undergone reinforcement learning based on a reward function such that the smaller the distance between the peg held by the virtual robot and the hole, the higher the reward. 4. The control device according to claim 3. (Appendix 5) the second trained model is a model that has undergone self-supervised learning based on data obtained by executing the simulation, and that, when the first information at time t, the second information at time t, and a candidate for the behavioral information at time t are input, outputs the first information at time t+1 and the behavioral information at time t+1; the determination unit inputs the first information at time t acquired by the first acquisition unit, the second information at time t acquired by the second acquisition unit, and candidates for the behavior information at time t to the second trained model, and determines behavior information of the real robot from the candidates for the behavior information at time t based on the second information at time t+1 output from the second trained model. 10. The control device according to claim 1 or 2. (Appendix 6) the task is a peg-in-hole task in which the real robot or the virtual robot inserts a peg held by the real robot or the virtual robot into a hole; the second information includes information representing a positional relationship between the peg and the hole; when determining the behavior information of the real robot from the candidates for the behavior information at time t, the determination unit determines the behavior information of the real robot from the candidates for the behavior information at time t based on information representing a positional relationship included in the second information. 6. The control device according to claim 5. (Appendix 7) a simulation unit that executes a simulation of a virtual robot on a computer, the virtual robot having some parts more flexible than other parts, performing a task involving contact; a learning data acquisition unit that acquires, from data obtained during execution of the simulation, first learning information representing an observation state of the virtual robot, second learning information that is not observed in a real space in which a real robot corresponding to the virtual robot exists, and learning behavior information representing behavior of the virtual robot; a learning unit that generates, using supervised machine learning based on the first information for learning, the second information for learning, and the behavioral information for learning acquired by the learning data acquisition unit, a first trained model that outputs the second information when the first information is input, and also generates, using machine learning, a second trained model for determining the behavior of the real robot; A trained model generation device including: (Appendix 8) acquiring first information representing an observed state of a real robot in real space, the real robot having a part more flexible than other parts, when the real robot performs a task involving contact; acquiring the second information by inputting the first information acquired by the acquisition unit into a first trained model that outputs second information that is not observed in the real space when the first information is input; determining behavior information of the real robot using a second trained model for determining the behavior of the real robot, the first information acquired by the first acquisition unit, and the second information acquired by the second acquisition unit; controlling the real robot so that the behavior represented by the determined behavior information is realized; the first trained model is a model obtained by supervised machine learning based on data obtained by executing a simulation in which a virtual robot, which is a robot on a computer corresponding to the real robot, performs the task; and The second trained model is a model that has been machine-learned based on data obtained by executing the simulation. A control method for computer-implemented processing. (Appendix 9) acquiring first information representing an observed state of a real robot in real space, the real robot having a part more flexible than other parts, when the real robot performs a task involving contact; acquiring the second information by inputting the first information acquired by the acquisition unit into a first trained model that outputs second information that is not observed in the real space when the first information is input; determining behavior information of the real robot using a second trained model for determining the behavior of the real robot, the first information acquired by the first acquisition unit, and the second information acquired by the second acquisition unit; controlling the real robot so that the behavior represented by the determined behavior information is realized; the first trained model is a model obtained by supervised machine learning based on data obtained by executing a simulation in which a virtual robot, which is a robot on a computer corresponding to the real robot, performs the task; and The second trained model is a model that has been machine-learned based on data obtained by executing the simulation. A control program that causes a computer to execute a process. (Appendix 10) A simulation is carried out on a computer in which a virtual robot, one part of which is more flexible than other parts, performs a task involving contact; acquiring, from data obtained during the execution of the simulation, first information for learning that represents an observation state of the virtual robot, second information for learning that is not observed in a real space in which a real robot corresponding to the virtual robot exists, and behavior information for learning that represents a behavior of the virtual robot; generating, using supervised machine learning based on the acquired first information for learning, second information for learning, and behavioral information for learning, a first trained model that outputs the second information when the first information is input, and generating, using machine learning, a second trained model for determining the behavior of the real robot; A method for generating trained models in which processing is performed by a computer. (Appendix 11) A simulation is carried out on a computer in which a virtual robot, one part of which is more flexible than other parts, performs a task involving contact; acquiring, from data obtained during the execution of the simulation, first information for learning that represents an observation state of the virtual robot, second information for learning that is not observed in a real space in which a real robot corresponding to the virtual robot exists, and behavior information for learning that represents a behavior of the virtual robot; generating, using supervised machine learning based on the acquired first information for learning, second information for learning, and behavioral information for learning, a first trained model that outputs the second information when the first information is input, and generating, using machine learning, a second trained model for determining the behavior of the real robot; A trained model generation program that allows a computer to execute processing. [Explanation of symbols]

[0108] 10. Control System 12. Robot 12A Arm 12G grip part 12S flexible part 14 Control device 18 Simulation Department 20 Learning data acquisition unit 22 Learning Department 24 Acquisition Department 26 Acquisition Department 28 Decision Section 30 Control Unit 32 Data storage unit 34 Trained model memory

Claims

1. a first acquisition unit that acquires first information representing an observation state of a real robot, the real robot being a robot in real space, where one part of the robot is more flexible than other parts, when the real robot performs a task involving contact; a second acquisition unit that acquires second information by inputting the first information acquired by the first acquisition unit into a first trained model that outputs second information that is not observed in the real space when the first information is input; a determination unit that determines behavior information of the real robot using a second learned model for determining the behavior of the real robot, the first information acquired by the first acquisition unit, and the second information acquired by the second acquisition unit; a control unit that controls the real robot so that the behavior represented by the behavior information determined by the determination unit is realized, the first trained model is a model obtained by supervised machine learning based on data obtained by executing a simulation in which a virtual robot, which is a robot on a computer corresponding to the real robot, performs the task; and The second trained model is a model that has been machine-learned based on data obtained by executing the simulation. Control device.

2. The task is a task that requires the robot itself or an instrument held by the robot to perform an operation while coming into contact with another object. The control device according to claim 1 .

3. the second trained model is a model that has undergone reinforcement learning based on data obtained by executing the simulation, and that outputs behavior information of the real robot when the first information and the second information are input; The determination unit acquires the behavioral information by inputting the first information acquired by the first acquisition unit and the second information acquired by the second acquisition unit to the second trained model. The control device according to claim 1 or 2.

4. the task is a peg-in-hole task in which the real robot or the virtual robot inserts a peg held by the real robot or the virtual robot into a hole; the second trained model is a model that has undergone reinforcement learning based on a reward function such that the smaller the distance between the peg held by the virtual robot and the hole, the higher the reward. The control device according to claim 3 .

5. the second trained model is a model that has undergone self-supervised learning based on data obtained by executing the simulation, and that, when the first information at time t, the second information at time t, and candidates for the behavioral information at time t are input, outputs the first information at time t+1 and the behavioral information at time t+1; the determination unit inputs the first information at time t acquired by the first acquisition unit, the second information at time t acquired by the second acquisition unit, and candidates for the behavior information at time t to the second trained model, and determines behavior information of the real robot from the candidates for the behavior information at time t based on the second information at time t+1 output from the second trained model. The control device according to claim 1 or 2.

6. the task is a peg-in-hole task in which the real robot or the virtual robot inserts a peg held by the real robot or the virtual robot into a hole; the second information includes information representing a positional relationship between the peg and the hole; the determination unit, when determining the behavior information of the real robot from the candidates for the behavior information at time t, determines the behavior information of the real robot from the candidates for the behavior information at time t based on information representing a positional relationship included in the second information. The control device according to claim 5 .

7. a simulation unit that executes a simulation of a virtual robot on a computer, the virtual robot having some parts more flexible than other parts, performing a task involving contact; a learning data acquisition unit that acquires, from data obtained during execution of the simulation, first learning information representing an observation state of the virtual robot, second learning information that is not observed in a real space in which a real robot corresponding to the virtual robot exists, and learning behavior information representing behavior of the virtual robot; a learning unit that generates, using supervised machine learning based on the first information for learning, the second information for learning, and the behavioral information for learning acquired by the learning data acquisition unit, a first trained model that outputs the second information when the first information is input, and also generates, using machine learning, a second trained model for determining the behavior of the real robot; A trained model generation device including:

8. acquiring first information representing an observation state of a real robot in real space, the real robot having a part more flexible than another part, when the real robot performs a task involving contact; acquiring the second information by inputting the acquired first information into a first trained model that outputs second information that is not observed in the real space when the first information is input; determining behavior information of the real robot using a second learned model for determining the behavior of the real robot, the acquired first information, and the acquired second information; controlling the real robot so that the behavior represented by the determined behavior information is realized; the first trained model is a model obtained by supervised machine learning based on data obtained by executing a simulation in which a virtual robot, which is a robot on a computer corresponding to the real robot, performs the task; and The second trained model is a model that has been machine-learned based on data obtained by executing the simulation. A control method for computer-implemented processing.

9. acquiring first information representing an observation state of a real robot in real space, the real robot having a part more flexible than another part, when the real robot performs a task involving contact; acquiring the second information by inputting the acquired first information into a first trained model that outputs second information that is not observed in the real space when the first information is input; determining behavior information of the real robot using a second learned model for determining a behavior of the real robot, the first information, and the acquired second information; controlling the real robot so that the behavior represented by the determined behavior information is realized; the first trained model is a model obtained by supervised machine learning based on data obtained by executing a simulation in which a virtual robot, which is a robot on a computer corresponding to the real robot, performs the task; and The second trained model is a model that has been machine-learned based on data obtained by executing the simulation. A control program that causes a computer to execute a process.

10. A simulation is carried out on a computer in which a virtual robot, one part of which is more flexible than other parts, performs a task involving contact; acquiring, from data obtained during the execution of the simulation, first information for learning that represents an observation state of the virtual robot, second information for learning that is not observed in a real space in which a real robot corresponding to the virtual robot exists, and behavior information for learning that represents a behavior of the virtual robot; generating, using supervised machine learning based on the acquired first information for learning, second information for learning, and behavior information for learning, a first trained model that outputs the second information when the first information is input, and generating, using machine learning, a second trained model for determining behavior of the real robot; A method for generating trained models in which processing is performed by a computer.

11. A simulation is carried out on a computer in which a virtual robot, one part of which is more flexible than other parts, performs a task involving contact; acquiring, from data obtained during the execution of the simulation, first information for learning that represents an observation state of the virtual robot, second information for learning that is not observed in a real space in which a real robot corresponding to the virtual robot exists, and behavior information for learning that represents a behavior of the virtual robot; generating, using supervised machine learning based on the acquired first information for learning, second information for learning, and behavior information for learning, a first trained model that outputs the second information when the first information is input, and generating, using machine learning, a second trained model for determining behavior of the real robot; A trained model generation program that allows a computer to execute processing.