Trained model generation device, control device, trained model generation method, and trained model generation program

The learned model generation device efficiently generates models for robot object movement tasks by using symmetry-based auxiliary loss functions, addressing the inefficiencies of existing techniques in handling extensive data requirements and high computational costs.

WO2025134930A1PCT designated stage expired Publication Date: 2025-06-26OMRON CORP
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
PCT/JP2024/044112
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2023-12-22
Filing Date
2024-12-12
Publication Date
2025-06-26

AI Technical Summary

Technical Problem

Existing techniques for teaching robots to grasp objects require extensive learning image data and are inefficient in generating learned models for tasks like moving objects, due to high computational costs and the need for large amounts of movement history data.

Method used

A learned model generation device and method that acquire pairs of movement history data from different starting points and use an auxiliary loss function to generate a learned model efficiently, by assuming the original and converted movement histories are the same, thereby reducing computational costs.

Benefits of technology

The approach allows for the efficient generation of learned models for controlling object movement, reducing computational costs and improving model accuracy by leveraging symmetry in movement data.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure JP2024044112_26062025_PF_FP_ABST
    Figure JP2024044112_26062025_PF_FP_ABST
Patent Text Reader

Abstract

This control device acquires a pair of sets of data, namely: first data representing movement history of a training subject from a first start point to an end point; and second data representing movement history of the training subject from a second start point to the end point. By adding, to a first loss function used when executing reinforcement learning, a second loss function representing the difference between an output value when the first data is inputted to a trained model and an output value when the second data is inputted to the trained model, the control device 14 sets an overall loss function for generating the trained model. On the basis of the acquired pair, the control device performs reinforcement learning on the trained model such that the output value of the overall loss function decreases, thereby generating a trained model for outputting movement data representing displacement of the position of a subject when position data for the subject is inputted.
Need to check novelty before this filing date? Find Prior Art

Description

Trained model generation device, control device, trained model generation method, and trained model generation program

[0001] The present disclosure relates to a trained model generation device, a control device, a trained model generation method, and a trained model generation program.

[0002] Conventionally, there is known a technique for teaching a robot how to grasp an object (see, for example, Reference 1: Xupeng Zhu, Dian Wang, Ondrej Biza, Guanang Su, Robin Walters, Robert Platt, "Sample Efficient Grasp Learning Using Equivariant Models", https: / / arxiv.org / abs / 2202.09468.). This technique increases the amount of learning image data containing an object, and the robot learns the movements required to grasp an object by moving its gripper based on the learning images.

[0003] The technology in the above-mentioned document 1 is a technology for increasing the amount of learning image data containing an object when a robot is made to learn an action for grasping an object. Furthermore, in the above-mentioned document 1, an action value function in reinforcement learning is learned using a convolutional neural network. A convolutional neural network is a model that can take into account the symmetry of an image containing an object. In the above-mentioned document 1, by realizing the action value function using a convolutional neural network, for example, when an object shown in an image rotates, the value of the action value function also changes in accordance with the rotation of the object.

[0004] On the other hand, for example, a robot may execute a task such as moving an object. In this case, when the robot executes reinforcement learning, the robot needs to learn the path of the object's movement. For example, when the robot executes a task called "peg-in-hole," which involves inserting a peg, which is an example of an object, into a hole, the robot needs to learn what path the peg should take to move it into the hole.

[0005] In the above-mentioned document 1, an action value function that takes into account the rotational symmetry of the object itself is realized by a convolutional neural network. Therefore, even if the technology in the above-mentioned document 1 is used, only a trained model that takes into account the rotational symmetry of the object itself can be obtained, and this trained model is a trained model used when grasping an object. Even if an attempt is made to obtain a trained model that is used when performing a task of moving an object using the technology in the above-mentioned document 1, it would require a huge computational cost, making it impossible to efficiently generate a trained model.

[0006] In order to generate a trained model to be used when executing a task of moving an object, it is necessary to use a huge amount of training data representing the movement history when the object is actually moved. However, when generating a trained model using such a huge amount of training data, a huge computational cost is required, which poses a problem that the trained model cannot be generated efficiently.

[0007] The present disclosure has been made in consideration of the above points, and aims to efficiently generate a trained model to be used when performing a task of moving an object from a start point to an end point.

[0008] In order to achieve the above object, the trained model generation device according to the present disclosure is a trained model generation device that generates a trained model for controlling the movement of an object, and includes: a training acquisition unit that acquires pairs of first data representing the movement history of a training object from a first start point to an end point and second data representing the movement history of the training object from a second start point to an end point; a setting unit that sets an overall loss function for generating the trained model by adding, to a first loss function used when performing reinforcement learning, a second loss function that represents the difference between the output value when the first data is input to the training model and the output value when the second data is input to the training model; and a learning unit that generates the trained model that outputs movement data representing the displacement of the position of the object when position data of the object is input, by performing reinforcement learning on the training model based on the pairs acquired by the training acquisition unit so that the output value of the overall loss function becomes smaller.

[0009] The trained model generation method of the present disclosure is a trained model generation method for generating a trained model for controlling the movement of an object, and includes the steps of: acquiring a pair of first data representing the movement history of a training object from a first start point to an end point and second data representing the movement history of the training object from a second start point to an end point; setting an overall loss function for generating the trained model by adding a second loss function representing the difference between the output value when the first data is input to the training model and the output value when the second data is input to the training model to a first loss function used when performing reinforcement learning; and reinforcing learning the training model based on the acquired pair so as to reduce the output value of the overall loss function, thereby generating the trained model that outputs movement data representing the displacement of the position of the object when position data of the object is input.

[0010] The trained model generation program of the present disclosure is a trained model generation program that generates a trained model for controlling the movement of an object, and includes the steps of: acquiring a pair of first data representing the movement history of a training object from a first starting point to an end point and second data representing the movement history of the training object from a second starting point to an end point; setting an overall loss function for generating the trained model by adding a second loss function representing the difference between the output value when the first data is input to the training model and the output value when the second data is input to the training model to a first loss function used when performing reinforcement learning; and reinforcing learning the training model based on the acquired pair so as to reduce the output value of the overall loss function, thereby generating the trained model that outputs movement data representing the displacement of the position of the object when position data of the object is input.

[0011] The trained model generation device, control device, trained model generation method, and trained model generation program disclosed herein can efficiently generate trained models to be used when performing a task of moving an object from a start point to an end point.

[0012] 1 is a diagram for explaining a control system of this embodiment; FIG. 2 is a diagram for explaining a tip portion of an arm; FIG. 3 is a diagram for explaining a trajectory when a peg is moved toward a hole; FIG. 4 is a diagram for explaining a trajectory when a peg is moved toward a hole; FIG. 5 is a diagram for explaining expansion of learning data; FIG. 6 is a diagram for explaining expansion of learning data taking symmetry into consideration; FIG. 7 is a diagram for explaining a learning model of this embodiment; FIG. 8 is a block diagram showing a schematic configuration of a control system of this embodiment; FIG. 9 is a block diagram showing a hardware configuration of a control device according to this embodiment; FIG. 10 is a flowchart showing the flow of a trained model generation process in this embodiment; FIG. 11 is a flowchart showing the flow of a control process in this embodiment; FIG. 12 is a diagram showing the results of this embodiment; FIG. 13 is a diagram showing the results of this embodiment;

[0013] An example of an embodiment of the present disclosure will be described below with reference to the drawings. In this embodiment, a control system equipped with a control device according to the present disclosure will be described as an example. Note that the same reference numerals are used in the drawings to designate identical or equivalent components and parts. Furthermore, the dimensions and proportions of the drawings are exaggerated for the sake of explanation and may differ from the actual proportions.

[0014] Fig. 1 is a diagram for explaining this embodiment. As shown in Fig. 1, a control system 10 of this embodiment includes a robot 12 and a control device 14. The robot 12 includes an arm 16. The robot 12 performs a peg-in-hole task, which is a task of inserting a peg T into a hole H, by operating the arm 16 in response to a control signal output from the control device 14.

[0015] FIG. 2 is a diagram illustrating the tip portion of the arm 16. As shown in FIG. 2, the tip portion of the arm 16 includes an arm body 16A and a gripper 16B. The arm body 16A and the gripper 16B are connected by a spring 16C, which is an example of an elastic body. The arm body 16A and the gripper 16B normally operate as a single unit. For example, when a peg T comes into contact with the floor surface F of a hole H, the arm body 16A and the gripper 16B are controlled to be connected by the spring 16C. As a result, for example, when the peg T comes into contact with the floor surface F, the arm body 16A and the gripper 16B are connected by the spring 16C, making it easier to search for the position of the hole H.

[0016] When detecting whether the peg T has come into contact with the floor surface F, for example, a force sensor (not shown) described below is installed on the arm 16, and when the force sensor detects a force due to the peg T coming into contact with the floor surface F, the arm body 16A and the gripper 16B are controlled to be connected by the spring 16C.

[0017] 3 and 4 are diagrams for explaining the trajectory when the peg T is moved toward the hole H. As shown in FIG. 3, the trajectories m1 and m2 when the peg T is moved toward the hole H are in a line-symmetric relationship. Specifically, the trajectory m2 in FIG. 3 is obtained by inverting the trajectory m1 in FIG. 3 with respect to the Y axis. Furthermore, as shown in FIG. 4, the trajectories m3, m4, m5, and m6 are in a point-symmetric relationship with respect to the hole H as the center point.

[0018] 3 and 4, a known reinforcement learning algorithm can be used to train the robot 12. When using reinforcement learning to generate a trained model for controlling the movement of the arm 16 of the robot 12, a large amount of training data is required.

[0019] Therefore, in this embodiment, the training data is increased by converting the actually obtained data. For example, a trajectory m2 is generated by inverting the trajectory m1 of the peg T shown in FIG. 3 with respect to the Y axis. Then, the movement history corresponding to the trajectory m1 and the movement history corresponding to the trajectory m2 are set as training data. Furthermore, for example, a trajectory m3 of the peg T shown in FIG. 4 is rotated around the hole H to generate trajectories m4, m5, and m6. Then, the movement history corresponding to each generated trajectory is set as training data.

[0020] However, when generating a trained model using the training data increased as described above, there is a problem in that the computational costs increase and it is not possible to generate a trained model efficiently because each piece of the increased huge amount of training data is different and separate data.

[0021] Therefore, in this embodiment, reinforcement learning is performed assuming that the original movement history and the converted movement history among the increased number of learning data are identical, thereby generating a trained model to be used by the robot 12 to move an object. Specifically, a new auxiliary loss function is added to the existing loss function used in reinforcement learning. This auxiliary loss function is a loss function that includes the difference between the output value when data representing the original movement history is input and the output value when data representing the converted movement history is input. By generating a trained model so that the output value of this auxiliary loss function is small, it is possible to efficiently generate a trained model to be used when executing a task of moving an object from a start point to an end point.

[0022] The specific details will be explained below.

[0023] <Framework of this embodiment> Conventionally, a known motion capture system has been used to sequentially detect the relative positional relationship between a robot's arm, a hole, and a peg, and to control the movement of the robot's arm. When a robot is made to execute a peg-in-hole task, a motion capture system is also used to control the movement of the robot's arm. When a motion capture system is used, it is possible to directly observe the relative positional relationship between the arm, the hole, and the peg.

[0024] In contrast to this, in this embodiment, a force sensor (not shown) is installed in the arm 16 of the robot 12, and the movement of the arm 16 of the robot 12 is controlled based on the force values ​​output from the force sensor. The values ​​output from the force sensor are three-axis force values ​​and three-axis moment values.

[0025] For this reason, this embodiment employs a partially observable Markov decision process (POMDP) ​​framework. The POMDP framework is used when only part of the state of a target system can be observed. In this embodiment, the movement of the arm 16 of the robot 12 is controlled based on force values ​​detected by a force sensor without directly observing the relative positional relationship between the arm 16, the hole H, and the peg T. Therefore, this POMDP framework is employed to control the movement of the arm 16 of the robot 12.

[0026] In addition, this embodiment employs a Deep Reinforcement Learning (DRL) approach and a Soft-Actor Critic (SAC) framework. SAC is disclosed in, for example, Reference 1 below.

[0027] Reference 1: T. Haarnoja, A. Zhou, P. Abbeel, and S. Levine, “Soft actor-critic: Off policy maximum entropy deep reinforcement learning with a stochastic actor,” in International Conference on Machine Learning, 2018, pp.1861-1870.

[0028] (Partially Observable Markov Decision Process) The framework of POMDP is defined by a tuple (S, A, Ω, T, R, O). S is the state space, A is the action space, and Ω is the observation space. T represents the dynamic function, R represents the reward value, and O is the observation function. After an agent takes an action a, state s changes to state s' according to the dynamic function T(s, a, s'). Note that when transitioning from state s to state s', an observation o determined by the observation function O(a, s', o) is output. In order for the agent to behave optimally, a movement history h, which is time-series data of the observation o and the action a, is required, as shown in the following equation: t It is necessary to determine appropriately.

[0029]

[0030] The movement history h is calculated so that the total reward J expressed by the following formula is maximized. t The strategy π(h t By generating R(s) in the following equation, the agent can take appropriate action a in response to the observation o. t , a t ) is the state s of the agent at time t. t Action A t is the reward value obtained when taking γ. γ is the discount rate. E represents the expected value.

[0031]

[0032] In this embodiment, the movement history h t The strategy π(h t ) and movement history h t Action value function Q(s t , a t ) is learned. Details will be explained later.

[0033] (Formulation) Fig. 5 is a diagram for explaining the expansion of the learning data. As shown in Fig. 5, when the XY coordinate system is set with the hole H as the origin, the starting point p 0 Assume that the coordinate data (x, y) of the above is actually obtained training data. In this case, the function that performs the transformation with the X axis as the inversion axis is F x Let F be the function that performs the transformation with the Y axis as the inverted axis. y Also, let R be the function that performs counterclockwise rotation around the origin. n Let's say.

[0034] Here, we define a transformation set G whose components are transformation elements g that represent transformations of the training data. We also define an element e that does not transform the input. For example, if the transformation element g is F x If F x = {e, s x}, and s x represents the inversion operation with the X axis as the inversion axis. x , then, as shown in FIG.x *p 0 = {e * p 0 , s x *p 0}={p 0 , p 1} Similarly, if the transformation element g is F y If F x = {e, s y}, and s y represents the inversion operation with the Y axis as the inversion axis. As shown in FIG. y *p 0 = {e * p 0 , s y *p 0}={p 0 , p 5}. Here, for ease of notation, F x *F y F xy The transformation element g is expressed as F xy If F xy *p 0 = {p 0 , p 1 , p 4 , p 5}.

[0035] In this embodiment, a function R that generates n plane rotations {0, 2π / n, 4π / n, . . . , 2π(n−1) / n} is used. n is defined. In this case, e corresponds to a rotation of 0 degrees. The transformation element g is n If R n coordinate data p 0 By acting on p 0 For example, R 4 *p 0 = {p 0 , p 2 , p 4 , p 6}. Furthermore, F xy and R 4 When combined with xy *R 4 *p 0 = {p 0 , p 1 , p 2 , p3 , p 4 , p 5 , p 6 , p 7}.

[0036] As described above, when these functions are combined to generate learning data, the result is as follows: Furthermore, the positional relationship between the respective learning data is as shown in FIG.

[0037] F x *p 0 = {p 0 , p 1 F y *p 0 = {p 0 , p 5 R 4 *p 0 = {p 0 , p 2 , p 4 , p 6 F x *F y *p 0 =F xy *p 0 = {p 0 , p 1 , p 4 , p 5 F xy *R 4 *p 0 = {p 0 , p 1 , p 2 , p 3 , p 4 , p 5 , p 6 , p 7}

[0038] In this embodiment, |G| is defined as the size of the transformation set G. For example, |R n |=n, |F x |=|F y |=2, |F xy |=4, |F xy *R 4 |=8.

[0039] (Variables Set in This Embodiment) In this embodiment, the observation o∈Ω is the three-dimensional position (t x , ty , t z ) and three-axis torque (τ x , τ y , τ z ) and three-axial force (f x , f y , f z ) and continuous action a = (δ x , δ y , δ z ) is data corresponding to the three-dimensional displacement of the arm body 16A. In order to utilize the flexibility of the arm 16 provided by the spring 16C, it is preferable to adjust the position of the arm 16 of the robot 12 after the peg T has contacted the floor surface F. The state s∈S is data including the relative posture between the peg T and the hole H, and s=p p2h In this embodiment, since the state s is not directly observed, a partially observable Markov decision process is adopted. The observation o, the action a, the state s, and the reward r are expressed by the following equation (1). As shown in equation (1), the reward r is 1 if the task is successful.

[0040]

[0041] (Proposed Method) Fig. 6 is a diagram for explaining the expansion of learning data taking symmetry into consideration. Fig. 6 shows the movement history of the arm 16 projected onto the XY plane. The movement history h shown in Fig. 6 T 0 is the actual movement history of the arm body 16A of the robot 12, and is the movement history that successfully inserts the peg T into the hole H. Here, the transformation set G=F xy is selected, and the movement history h T 0 Against F xy When the action of the three movement histories (h ’ T 1 , h ’ T 2 , h ’ T 3 ) is generated. T 0 and three movement histories (h ’ T1 , h ’ T 2 , h ’ T 3 The relationship between these three movement histories (h ’ T 1 , h ’ T 2 , h ’ T 3 ) is used as learning data taking into account symmetry. T 0 corresponds to the movement history of the arm 16 of the robot 12 when the task is successful. The movement history when the task is unsuccessful can also be used as learning data.

[0042] Also, Fig. 6 shows the symmetry of the action value function Q. The current movement history h t As the movement history h T 0 Given Q(h' t , a' t ) and Q(h' t 1 , a' t 1 ), Q(h' t 2 , a' t 2 ), and Q(h′ t 3 , a' t 3 ) must be equal to a' t i (i=1, 2, 3) is a transformed by the transformation set G t Therefore, it is necessary to generate an action-value function Q that is invariant to the transformation by the transformation set G, as shown in the following formula (2).

[0043]

[0044] Also, action a at time t+1 t+1 = π(h t ) and the movement history h tWhen the movement history converted from t+1 Similarly to the conversion of , it is possible to derive the following equation (3).

[0045]

[0046] Therefore, the transformation set G=F xy and given the movement history (h ’ T 1 , h ’ T 2 , h ’ T 3 The action at time t+1 corresponding to (h ’ t+1 1 , h ’ t+1 2 , h ’ t+1 3 ) Action a shown in FIG. t+1 , a ’ t+1 1 , a ’ t+1 2 , and a ’ t+1 3 is represented by an arrow, and it is indicated that the arm 16 of the robot 12 is displaced in the direction of the arrow.

[0047] The behavior data a before conversion shown in FIG. t+1 The output policy π and the converted action data a ’ t+1 1 Similarly, the policy π that outputs the pre-conversion behavior data a t+1 The output policy π and the converted action data a ’ t+1 2 The policy π that outputs the action data a before conversion must be equal. t+1 The output policy π and the converted action data a ’ t+1 3 The strategies π that output must be equal.

[0048] (Data Expansion Considering Symmetry) When a transformation element gεG is given, a data expansion equation such as the following equation (4) is defined.

[0049]

[0050] The movement history h is calculated by the transformation element g. t When the movement history h t A transformation factor g is applied to every observation o and every action a in . Here, the following definitions are made:

[0051] (Definition 1) g*o: p shown in FIG. 0 Similarly, F xy is the Z-axis element (t z , τ z , f z ) is not changed, and (t x , t y ), (τ x , τ y ), (f x , f y ) (Definition 2) g*a: g*o, and F xy is the Z-axis element a z In addition, the p shown in FIG. 0 Similarly, F xy is (t x , t y ), (τ x , τ y ), (f x , f y ) only.

[0052] (Loss Function Taking Transformation into Account) In this embodiment, the action-value function Q and the policy π are learned so that the above formulas (2) and (3) are satisfied. As described above, in this embodiment, the action-value function Q and the policy π are learned using SAC disclosed in the above-mentioned Reference 1.

[0053] 7 is a diagram showing a recurrent SAC model in this embodiment. The recurrent SAC model is disclosed in the following reference 2.

[0054] Reference 2: T. Ni, B. Eysenbach, and R. Salakhutdinov, “Recurrent model-free RL can be a strong baseline for many POMDPs,” in International Conference on Machine Learning, 2022, pp. 16691-16723.

[0055] "ACTOR" in Fig. 7 is a learning model (hereinafter simply referred to as "actor") that learns the policy π, and "CRITICS" in Fig. 7 is a learning model (hereinafter simply referred to as "critic") that learns the action-value function Q. "RNN" in Fig. 7 represents a recurrent neural network, "MLP" represents a multilayer perceptron, and "Embedder" is a neural network model for embedding observation o and action a.

[0056] There are two differences between the recurrent SAC model of this embodiment and the SAC model disclosed in the above-mentioned Reference 2. The first difference is that the training data is increased by performing data augmentation as described above. The second difference is that an auxiliary loss function is introduced so that the above-mentioned formulas (2) and (3) are satisfied. The auxiliary function will be explained below. Note that the auxiliary loss function L sym C (Q) and auxiliary loss function L sym a (π) is an example of the second loss function of the present disclosure.

[0057] (Auxiliary Loss Function for Critic) As shown in the above formula (2), in this embodiment, the critic needs to be trained so that the output value of the action value function Q for (h, a) before transformation is equal to the output value of the action value function Q for (g*h, g*a) transformed by g included in the transformation set G. For this reason, in this embodiment, the critic is trained so that the difference between the action value function Q(h, a) before transformation and the action value function Q(g*h, g*a) after transformation becomes small. Specifically, the auxiliary loss function L shown in the following formula (5) symC The critic is trained so that (Q) is minimized.

[0058]

[0059] As shown in FIG. 7, when the model disclosed in Reference 2 is adopted, two types of action value functions Q 1 Output value and action value function Q 2 When the SAC model in Reference 2 is not adopted, the critic outputs a single value of the action-value function Q.

[0060] (Auxiliary loss function for actors) In SAC, actors take action a and achieve the average μ a and the standard deviation of action a, σ a Therefore, in this embodiment, the average μ of the action a before conversion is output. h a and the average μ of the converted action a g*h a Specifically, the actor is trained so that the difference between the sym a The actor is trained to minimize (π).

[0061]

[0062] In addition, the average μ a and the standard deviation of action a, σ a The details of outputting the above are disclosed in Reference 3 below.

[0063] Reference 3: H. Nguyen, A. Baisero, D. Wang, C. Amato, and R. Platt, “Leveraging fully observable policies for learning under partial observability,” in Conference on Robot Learning, 2022.

[0064] (Control System 10) Fig. 8 is a block diagram showing a schematic configuration of the control system 10 of this embodiment. As shown in Fig. 8, the control system 10 includes a force sensor 11, a robot 12, and a control device 14. The control device 14 generates a trained model for controlling the movement of the arm 16 of the robot 12. The control device 14 also controls the movement of the arm 16 of the robot 12 using the generated trained model.

[0065] The force sensor 11 is attached to the arm 16 of the robot 12 and detects a three-axis torque (τ x , τ y , τ z ) and three-axial force (f x , f y , f z ) and the force sensor 11 outputs the obtained force values ​​to the control device 14.

[0066] In addition, among the observations o shown in the above formula (1), the three-dimensional position (t x , t y , t z ) can be easily calculated from the displacement represented by the action a. Therefore, the force sensor value (τ x , τ y , τ z , f x , f y , f z ) and the three-dimensional position (t x , t y , t z ) sets the observation o shown in equation (1) above.

[0067] The robot 12 is a robot as shown in FIG. 1, and performs a peg-in-hole task of inserting a peg T into a hole H by operating an arm 16.

[0068] 9 is a block diagram showing the hardware configuration of the control device 14 according to this embodiment. As shown in FIG. 9, the control device 14 includes a CPU (Central Processing Unit) 42, a memory 44, a storage device 46, an input / output I / F (Interface) 48, a storage medium reader 50, and a communication I / F 52. Each component is connected to each other via a bus 54 so as to be able to communicate with each other.

[0069] The storage device 46 stores a trained model generation program and a control program for executing each process described below. The CPU 42 is a central processing unit that executes various programs and controls each component. That is, the CPU 42 reads the programs from the storage device 46 and executes the programs using the memory 44 as a work area. The CPU 42 controls each component and performs various arithmetic processes in accordance with the programs stored in the storage device 46.

[0070] The memory 44 is configured with a RAM (Random Access Memory) and serves as a working area to temporarily store programs and data. The storage device 46 is configured with a ROM (Read Only Memory), an HDD (Hard Disk Drive), an SSD (Solid State Drive), etc., and stores various programs including the operating system and various data.

[0071] The input / output I / F 48 is an interface for inputting data from an external device and outputting data to an external device. Input devices for various inputs, such as a keyboard or a mouse, and output devices for various information outputs, such as a display or a printer, may also be connected. A touch panel display may be used as the output device, allowing it to function as an input device.

[0072] The storage medium reader 50 reads data stored in various storage media such as CD (Compact Disc)-ROM, DVD (Digital Versatile Disc)-ROM, Blu-ray Disc, and USB (Universal Serial Bus) memory, and writes data to the storage media.

[0073] The communication I / F 52 is an interface for communicating with other devices, and uses standards such as Ethernet (registered trademark), FDDI, and Wi-Fi (registered trademark).

[0074] Next, the functional configuration of the control device 14 will be described. As shown in Fig. 8, the control device 14 functionally includes a learning acquisition unit 17, a setting unit 18A, a learning unit 18B, an acquisition unit 20, and a control unit 22. A predetermined storage area of ​​the control device 14 is also provided with a data storage unit 24 and a trained model storage unit 26. Each functional configuration is realized by the CPU 42 reading each program stored in the storage device 46, expanding it into the memory 44, and executing it.

[0075] The data storage unit 24 stores force values ​​detected by the force sensor 11. The data storage unit 24 also stores control data when the arm 16 of the robot 12 operates. For example, the action (δ x , δ y , δ z ) is stored.

[0076] The learned model storage unit 26 stores the learned policy π and the learned action value function Q generated by the processing described below.

[0077] First, the learning acquisition unit 17 and the learning unit 18B generate a learned policy π, which is a learned model for controlling the movement of the arm 16 of the robot 12.

[0078] The movement history h of the arm 16 of the robot 12 when performing the peg-in-hole task tVarious data including the above are stored in the data storage unit 24. The movement history data at this time includes data when the peg-in-hole task is successful and data when the peg-in-hole task is unsuccessful. Note that, hereinafter, data representing the movement history obtained by the actual operation of the arm 16 of the robot 12 will be simply referred to as "first data". Note that the first data is data representing the movement history obtained by the actual operation of the arm 16 of the robot 12 from the first starting point p 0 This is the movement history when the robot is moved from the hole H to the end point.

[0079] The learning acquisition unit 17 reads out a plurality of first data from the data storage unit 24. Then, the learning acquisition unit 17 applies a transformation element g included in the transformation set G to each of the plurality of first data to generate a plurality of second data representing a movement history that is symmetrical with the first data. As shown in Fig. 5 and Fig. 6, the second data is generated by moving the arm main body 16A, which is an example of an object, from a second starting point p 1 , p 2 , p 3 , p 4 , p 5 , p 6 , p 7 The movement history is the movement history when the robot is moved from the first data to the end point, hole H. In this way, the learning data acquisition unit 17 acquires a plurality of pairs of the first data and the second data. This increases the amount of learning data, making it possible to generate a trained model with high accuracy.

[0080] The setting unit 18A is configured to set a known loss function L A , the auxiliary loss function L shown in the above equation (5) sym C By adding (Q), the overall loss function L for generating a trained model is obtained. A C As shown in the above equation (5), the auxiliary loss function L sym C (Q) is a loss function that represents the difference between the output value when the first data is input to the critic and the output value when the second data is input to the critic. For example, the loss function L A and the auxiliary loss function L sym CThe sum of (Q) and the overall loss function L A C is set as

[0081] In addition, the setting unit 18A sets a known loss function L B , the auxiliary loss function L shown in the above equation (6) sym a By adding (π), the overall loss function L for generating the trained model is obtained. B a As shown in the above equation (6), the auxiliary loss function L sym a (π) is a loss function that represents the difference between the output value when the first data is input to the actor and the output value when the second data is input to the actor. For example, the loss function L B and the auxiliary loss function L sym a (π) is the overall loss function L B a is set as

[0082] Note that the known loss function L A and L B is a loss function used in, for example, the SAC algorithm. A and L B is an example of the first loss function of the present disclosure.

[0083] The learning unit 18B learns the above-mentioned actor and critic by performing reinforcement learning based on the multiple pairs acquired by the learning acquisition unit 17. Note that when performing reinforcement learning, the learning unit 18B performs reinforcement learning on the learning model, assuming that the first data and the second data in the pair are identical. As a result, a learned policy π corresponding to the actor and an action-value function Q corresponding to the critic are generated.

[0084] The learned policy π is the three-dimensional position (t x , t y , t z ) is input, the displacement (δ x , δ y , δ z) The learned policy π is used to operate the arm 16 of the robot 12.

[0085] When performing reinforcement learning, the learning unit 18B uses the overall loss function L set by the setting unit 18A. A C When performing reinforcement learning, the learning unit 18B performs reinforcement learning so that the overall loss function L set by the setting unit 18A is reduced. B a A learned policy is generated by performing reinforcement learning so that π is small. As described above, the policy π corresponds to the actor, and the action value function Q corresponds to the critic.

[0086] Then, the learning unit 18B stores the learned policy π corresponding to the actor and the learned action-value function Q corresponding to the critic in the learned model storage unit 26.

[0087] Once the learned policy π is stored in the learned model storage unit 26, it becomes possible to use the learned policy π to control the movement of the arm 16 of the robot 12. Therefore, the acquisition unit 20 and the control unit 22 use the learned policy π stored in the learned model storage unit 26 to control the movement of the arm 16 of the robot 12.

[0088] The acquisition unit 20 acquires the three-dimensional position (t x , t y , t z ) is obtained.

[0089] The control unit 22 reads out the learned policy π stored in the learned model storage unit 26. Then, the control unit 22 reads out the three-dimensional position (t x , t y , t z ) is input to the trained policy π. The trained policy π provides the displacement (δ x , δ y , δ z) is output. The control unit 22 acquires the movement data output from the learned policy π and controls the position of the arm 16 based on the movement data. Specifically, the control unit 22 calculates the displacement (δ x , δ y , δ z ) is output to the arm 16 of the robot 12.

[0090] Next, the operation of the control system 10 according to this embodiment will be described.

[0091] First, the movement history of the arm 16 of the robot 12 while the arm 16 is performing the peg-in-hole task is collected and input to the control device 14. Then, data related to the movement history is stored in the data storage unit 24. Then, when the control device 14 receives a predetermined instruction signal, the CPU 42 of the control device 14 reads out the trained model generation program from the storage device 46, loads it into the memory 44, and executes it. As a result, the CPU 42 functions as each functional component of the control device 14, and the trained model generation process shown in FIG. 10 is executed.

[0092] In step S100, the learning acquisition unit 17 reads out a plurality of first data stored in the data storage unit 24, thereby acquiring a plurality of first data.

[0093] In step S102, the setting unit 18A generates each piece of second data representing a movement history that is symmetrical with the first data by applying the transformation element g included in the transformation set G to each of the plurality of first data acquired in step S100. In this way, the setting unit 18A acquires a plurality of pairs of first data and second data.

[0094] In step S103, the setting unit 18A sets a known loss function L A , the auxiliary loss function L shown in the above equation (5) sym C By adding (Q), the overall loss function L for generating a trained model is obtained. A CIn step S103, a known loss function L used when performing reinforcement learning is set. B , the auxiliary loss function L shown in the above equation (6) sym a By adding (π), the overall loss function L for generating the trained model is obtained. B a Set.

[0095] In step S104, the learning unit 18B performs reinforcement learning to learn a critic, which is an action value function Q, and an actor, which is a policy π. Specifically, the learning unit 18B learns a total loss function L according to a known SAC algorithm. A C The learning unit 18B performs reinforcement learning on the critic, which is an action value function, so that the overall loss function L is reduced according to the known SAC algorithm. B a The actor, which is the policy, is reinforced and learned so that is small.

[0096] In step S106, the learning unit 18B stores the learned action-value function Q and the learned policy π obtained in step S104 in the learned model storage unit 26.

[0097] Next, when the control device 14 receives a predetermined instruction signal, the control device 14 executes the control process shown in FIG.

[0098] In step S200, the acquisition unit 20 acquires the three-dimensional position (t x , t y , t z ) at time t t = (t x , t y , t z , τ x , τ y , τ z , f x , f y , f z ) and action a at time t-1 t-1 = (δ x , δ y , δ z ) to obtain the

[0099] In step S202, the control unit 22 reads out the learned policy π stored in the learned model storage unit 26. Then, in step S202, the control unit 22 reads out the observation o at time t acquired in step S200. t and action a at time t-1 t-1 and are input to the learned policy π. The learned policy π provides the displacement of the position of the arm 16 (δ x , δ y , δ z ) is output. Therefore, the control unit 22 acquires the movement data output from the learned policy π.

[0100] In step S204, the control unit 22 controls the position of the arm 16 based on the movement data acquired in step S202. Specifically, the control unit 22 calculates the displacement (δ x , δ y , δ z ) is output to the arm 16 of the robot 12.

[0101] As described above, the control device according to this embodiment generates a trained model for controlling the movement of an object. The control device acquires a pair of first data representing the movement history of a training object from a first start point to an end point and second data representing the movement history of the training object from a second start point to an end point. The control device also sets an overall loss function for generating the trained model by adding an auxiliary loss function representing the difference between the output value when the first data is input to the training model and the output value when the second data is input to the training model to a known loss function used when performing reinforcement learning. When performing reinforcement learning based on the acquired pair, the control device reinforces learning the training model so as to reduce the output value of the overall loss function, thereby generating a trained model that outputs movement data representing the displacement of the object's position when position data of the object is input. This makes it possible to efficiently generate a trained model for use in performing a task of moving an object from a start point to an end point.

[0102] Specifically, the control device of this embodiment generates a trained model for controlling the movement of a robot arm, which is an example of an object. The first data and second data for controlling the movement of the robot arm are data representing the displacement of the arm's position and a force sense value output from a force sensor installed on the arm. The second data is data obtained by converting the first data and is data that is symmetrical with the first data. The control device generates a trained model that outputs arm movement data when arm position data is input.

[0103] According to the control device of this embodiment, reinforcement learning is performed assuming that the first data representing the movement history of the arm from the first starting point to the end point and the second data representing the movement history of the arm from the second starting point to the end point are identical, thereby reducing the computational cost when generating a trained model, and enabling the trained model to be generated efficiently.

[0104] Furthermore, according to the control device of this embodiment, by generating second data by converting first data and using the first data and second data as training data, it is possible to generate a trained model with higher accuracy.

[0105] Furthermore, according to the control device of this embodiment, observation o=(t x , t y , t z , τ x , τ y , τ z , f x , f y , f z ), and information about the relative posture between the peg T and the hole H is not required. Therefore, as long as it is possible to acquire the force sensor value obtained from the force sensor 11 and the position of the arm main body 16A, it is possible to perform a task of moving an object, such as a peg-in-hole task, without using motion capture technology or the like.

[0106] Next, an example will be described. In this example, a simulation was performed to verify the effectiveness of the proposed method. In this simulation, simulations of peg-in-holes were performed for hole shapes of triangle, square, pentagon, hexagon, and circle.

[0107] FIG. 12 shows the simulation results of this embodiment. "SAC-State" shown in FIG. 12 is the result when a policy is learned based on state s. Note that state s here refers to information including the relative positional relationship between the peg and the hole. "SAC-Obs" is the result when a policy is learned using only observation o. "RSAC" corresponds to "SAC-Obs" when a recurrent SAC model is used. These three agents, "SAC-State," "SAC-Obs," and "RSAC," are learned without utilizing data symmetry and data augmentation.

[0108] The method proposed in this embodiment corresponds to "RSAC-Aug-Aux." Note that "RSAC-Equi" corresponds to the method disclosed in the following Reference 4. The method in the following Reference 4 is a POMDP-based method.

[0109] Reference 4: H. Nguyen, A. Baisero, D. Klee, D. Wang, R. Platt, and C. Amato, “Equivariant reinforcement learning under partial observability,” in Conference on Robot Learning, 2023. [Online]. Available: https: / / openreview.net / forum?id=AnDDMQgM7-

[0110] The horizontal axis of Fig. 12, "Environment Step," represents the number of learning attempts, and "Evaluation Success Rate" represents the success rate for six peg-in-hole trials. As shown in Fig. 12, it can be seen that, regardless of the shape of the hole, "RSAC-Aug-Aux" has the highest success rate among the methods proposed in this embodiment.

[0111] 13 compares the results when both data symmetry and data extension are performed with the results when only one of them is performed. "Aug-Aux" represents the results when both are performed, "Aug" represents the results when only data extension is performed, and "Aux" represents the results when only data symmetry is used.

[0112] As can be seen from FIGS. 12 and 13, by using the method of this embodiment, it is possible to accurately and efficiently generate a learned policy to be used when moving an object.

[0113] In the above embodiment, the learning object and the object are described as pegs T or arm bodies 16A, but the present invention is not limited to this. Any object that can be moved may be used. For example, the method of this embodiment is not limited to peg-in-hole tasks, but can also be applied to picking tasks, navigation tasks, etc. Furthermore, the method of this embodiment can also be applied to tasks such as moving an object by autonomous travel. Furthermore, this embodiment may be applied to other movement control, such as control of the operation of parts other than the robot's arms, or control of autonomous travel, autonomous flight, or autonomous navigation of a mobile robot.

[0114] Furthermore, in the above embodiment, data extension is performed to generate second data from first data, but this is not limiting. Data extension does not have to be performed. For example, if it is appropriate to treat certain data A and other data B among multiple data on the same level in the action value function Q and the policy π, the certain data A and other data B may be treated as identical when performing reinforcement learning. Also, for example, the second data does not have to be data symmetrical with the first data. For example, the second data may be data obtained by performing some kind of transformation process on the first data. For example, the second data may be data obtained by temporally transforming the first data. Also, for example, in the case of a video, the second data may be data obtained by playing the video in reverse.

[0115] In the above embodiment, the action-value function Q and the policy π are learned according to the SAC algorithm, but the present invention is not limited to this. The action-value function Q and the policy π may be learned according to other reinforcement learning algorithms.

[0116] In the above embodiment, the movement history h is data including the action a and the state o, but is not limited to this. The movement history h can be changed as appropriate, and may include, for example, only the position of the target object.

[0117] In the above embodiment, the first data and the second data are the displacement of the arm position and the force sense value output from the force sensor attached to the arm, but the present invention is not limited to this. For example, the first data and the second data may be data including either the displacement of the arm position or the force sense value output from the force sensor attached to the arm.

[0118] In the above embodiment, the control device 14 executes both the trained model generation process of Fig. 10 and the control process of Fig. 11 , but this is not limiting. For example, a trained model generation device implemented by a computer separate from the control device 14 may be provided, and the trained model generation device may execute the trained model generation process of Fig. 10 , while the control device 14 executes the control process of Fig. 11 . In this case, the trained model generation device includes at least the learning acquisition unit 17, the setting unit 18A, and the learning unit 18B.

[0119] Furthermore, various processors other than the CPU may execute the processes executed by loading software (programs) in the above embodiments. Examples of processors in this case include programmable logic devices (PLDs) (such as field-programmable gate arrays (FPGAs)) whose circuit configuration can be changed after manufacture, and dedicated electrical circuits, such as application-specific integrated circuits (ASICs), which are processors having circuit configurations designed specifically to execute specific processes. Each process may be executed by one of these various processors, or by a combination of two or more processors of the same or different types (e.g., multiple FPGAs, or a combination of a CPU and an FPGA). The hardware structure of these various processors is, more specifically, an electrical circuit that combines circuit elements such as semiconductor elements.

[0120] In the above embodiment, the programs are pre-stored (installed) in a storage device, but this is not limiting. The programs may be provided in a form stored in a storage medium such as a CD-ROM, DVD-ROM, Blu-ray Disc, or USB memory. The programs may also be downloaded from an external device via a network.

[0121] (Supplementary Notes) The following supplementary notes are provided regarding aspects of the present disclosure.

[0122] (Supplementary Note 1) A trained model generation device that generates a trained model for controlling movement of an object, comprising: a training acquisition unit that acquires pairs of first data representing a movement history of a training object from a first start point to an end point and second data representing a movement history of the training object from a second start point to an end point; a setting unit that sets an overall loss function for generating the trained model by adding a second loss function that represents the difference between an output value when the first data is input to the training model and an output value when the second data is input to the training model to a first loss function used when performing reinforcement learning; and a learning unit that generates the trained model that outputs movement data representing a displacement of the position of the object when position data of the object is input, by performing reinforcement learning on the training model based on the pairs acquired by the training acquisition unit so that the output value of the overall loss function becomes smaller.

[0123] (Supplementary Note 2) The trained model generation device according to Supplementary Note 1, wherein the second data is data obtained by converting the first data. The training object is a robot arm, the first data and the second data are data representing a positional displacement of the arm and a force sense value output from a force sensor installed on the arm, the second data is data obtained by converting the first data and is data that is symmetrical with the first data, and the learning unit generates the trained model in which, when positional data of the arm is input, the movement data of the arm is output.

[0124] (Supplementary Note 3) The trained model generation device according to Supplementary Note 1 or Supplementary Note 2, wherein the training object is a robot arm, the first data and the second data are data including at least one of a positional displacement of the arm and a force value output from a force sensor installed on the arm, the second data is data obtained by converting the first data and is data that is symmetrical with the first data, and the learning unit generates the trained model in which, when positional data of the arm is input, movement data of the arm is output.

[0125] (Supplementary Note 4) The trained model generation device according to any one of Supplementary Notes 1 to 3, wherein a relationship between the movement history represented by the first data and the movement history represented by the second data is a line-symmetric or point-symmetric relationship.

[0126] (Supplementary Note 5) The trained model generation device according to any one of Supplementary Notes 1 to 4, wherein the second loss function is a loss function including a difference between an output value of an action value function when the first data is input to the action value function in reinforcement learning and an output value of the action value function when the second data is input to the action value function.

[0127] (Supplementary Note 6) The trained model generation device according to any one of Supplementary Notes 1 to 5, wherein the training model and the trained model correspond to policies in reinforcement learning, and the second loss function is a loss function including a difference between an output value of the policy when the first data is input to the policy and an output value of the policy when the second data is input to the policy.

[0128] (Supplementary Note 7) The trained model generation device according to any one of Supplementary Notes 1 to 6, wherein a policy in reinforcement learning is an actor in a Soft Actor-Critic algorithm, an action value function in reinforcement learning is a critic in a Soft Actor-Critic algorithm, the second loss function includes a difference between an output value of the critic when the first data is input to the critic and an output value of the actor when the second data is input to the critic, and a difference between an output value of the actor when the first data is input to the actor and an output value of the actor when the second data is input to the actor, and the learning unit, when performing reinforcement learning according to the Soft Actor-Critic algorithm, learns the action value function corresponding to the critic and learns the policy corresponding to the actor so that the overall loss function is small, and generates the trained model corresponding to the actor.

[0129] (Supplementary Note 8) A control device comprising: an acquisition unit that acquires position data of an object; and a control unit that acquires movement data of the object by inputting the position data of the object acquired by the acquisition unit into the trained model generated by the trained model generation device described in any one of Supplementary Notes 1 to 7, and controls the position of the object based on the movement data of the object.

[0130] (Supplementary Note 9) A trained model generation method for generating a trained model for controlling the movement of an object, the trained model generation method comprising the steps of: setting an overall loss function for generating the trained model by adding a second loss function representing the difference between the output value when the first data is input to the training model and the output value when the second data is input to the training model to a first loss function used when performing reinforcement learning; and generating the trained model that outputs movement data representing the displacement of the position of the object when position data of the object is input by performing reinforcement learning on the training model based on the acquired pairs so that the output value of the overall loss function becomes smaller.

[0131] (Supplementary Note 10) A trained model generation program for causing a computer to execute a process of: setting an overall loss function for generating the trained model by adding, to a first loss function used when performing reinforcement learning, a second loss function representing the difference between the output value when the first data is input to the training model and the output value when the second data is input to the training model; and generating the trained model that outputs movement data representing the displacement of the position of the object when position data of the object is input by performing reinforcement learning on the training model based on the acquired pairs so that the output value of the overall loss function becomes smaller.

[0132] The disclosure of Japanese Patent Application No. 2023-217347, filed on December 22, 2023, is incorporated herein by reference in its entirety. All documents, patent applications, and technical standards mentioned herein are incorporated herein by reference to the same extent as if each individual document, patent application, and technical standard was specifically and individually indicated to be incorporated by reference.

Claims

1. A trained model generation device that generates a trained model for controlling the movement of an object, comprising: a training acquisition unit that acquires pairs of first data representing a movement history of a training object from a first start point to an end point and second data representing the movement history of the training object from a second start point to an end point; a setting unit that sets an overall loss function for generating the trained model by adding a second loss function representing the difference between an output value when the first data is input to the training model and an output value when the second data is input to the training model to a first loss function used when performing reinforcement learning; and a learning unit that generates the trained model that outputs movement data representing a displacement of the position of the object when position data of the object is input, by performing reinforcement learning on the training model based on the pairs acquired by the training acquisition unit so that the output value of the overall loss function becomes smaller.

2. The trained model generation device according to claim 1, wherein the second data is data obtained by converting the first data.

3. The trained model generating device according to claim 1 or claim 2, wherein the learning object is a robot arm, the first data and the second data are data including at least one of a positional displacement of the arm and a force sensor value output from a force sensor installed on the arm, the second data is data obtained by converting the first data and is symmetrical with the first data, and the learning unit generates the trained model in which, when positional data of the arm is input, movement data of the arm is output.

4. The trained model generating device according to claim 1 or claim 2, wherein the relationship between the movement history represented by the first data and the movement history represented by the second data is a line-symmetric or point-symmetric relationship.

5. The trained model generation device according to claim 1 or claim 2, wherein the second loss function is a loss function including a difference between an output value of an action value function when the first data is input to the action value function in reinforcement learning and an output value of the action value function when the second data is input to the action value function.

6. A trained model generation device as described in claim 1 or claim 2, wherein the learning model and the trained model correspond to a policy in reinforcement learning, and the second loss function is a loss function including a difference between an output value of the policy when the first data is input to the policy and an output value of the policy when the second data is input to the policy.

7. The trained model generation device according to claim 1 or claim 2, wherein: a policy in reinforcement learning is an actor in a Soft Actor-Critic algorithm; an action value function in reinforcement learning is a critic in a Soft Actor-Critic algorithm; the second loss function includes a difference between an output value of the critic when the first data is input to the critic and an output value of the actor when the second data is input to the critic, and a difference between an output value of the actor when the first data is input to the actor and an output value of the actor when the second data is input to the actor; and the learning unit, when performing reinforcement learning according to the Soft Actor-Critic algorithm, learns the action value function corresponding to the critic and learns the policy corresponding to the actor so that the overall loss function is small, and generates the trained model corresponding to the actor.

8. A control device comprising: an acquisition unit that acquires position data of an object; and a control unit that acquires movement data of the object by inputting the position data of the object acquired by the acquisition unit into the trained model generated by the trained model generation device described in claim 1 or claim 2, and controls the position of the object based on the movement data of the object.

9. A trained model generation method for generating a trained model for controlling the movement of an object, the trained model generation method comprising the steps of: acquiring pairs of first data representing the movement history of a training object from a first start point to an end point and second data representing the movement history of the training object from a second start point to an end point; setting an overall loss function for generating the trained model by adding a second loss function representing the difference between an output value when the first data is input to the training model and an output value when the second data is input to the training model to a first loss function used when performing reinforcement learning; and performing reinforcement learning on the training model based on the acquired pairs so as to reduce the output value of the overall loss function, thereby generating the trained model that outputs movement data representing a displacement of the position of the object when position data of the object is input.

10. A trained model generation program for generating a trained model for controlling the movement of an object, comprising: acquiring pairs of first data representing the movement history of a training object from a first starting point to an end point and second data representing the movement history of the training object from a second starting point to an end point; setting an overall loss function for generating the trained model by adding a second loss function representing the difference between an output value when the first data is input to the training model and an output value when the second data is input to the training model to a first loss function used when performing reinforcement learning; and performing reinforcement learning on the training model based on the acquired pairs so as to reduce the output value of the overall loss function, thereby generating the trained model that outputs movement data representing the displacement of the position of the object when position data of the object is input.

Citation Information

Patent Citations

  • Execution of a peg-in-hole task with unknown slope

    JP2021531177A