Trained model generation device, control device, trained model generation method, and trained model generation program
Patent Information
- Application Number
- JP2023217347
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2023-12-22
- Publication Date
- 2025-07-03
Smart Images

Figure 2025100170000001_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to a learned model generation device, a control device, a learned model generation method, and a learned model generation program.
Background Art
[0002] Conventionally, a technique for teaching a robot an operation of grasping an object has been known (see, for example, Non-Patent Document 1). In this technique, learning image data in which an object appears is increased, and the robot learns an operation when moving its own gripper to grasp the object based on the learning image.
Prior Art Documents
Non-Patent Documents
[0003]
Non-Patent Document 1
Summary of the Invention
Problems to be Solved by the Invention
[0004] The technique of Non-Patent Document 1 is a technique for increasing learning image data in which an object appears when teaching a robot an operation for gripping an object. Further, in Non-Patent Document 1, a behavior value function in reinforcement learning is learned using a convolutional neural network. The convolutional neural network is a model capable of considering the symmetry of an image in which an object appears. In Non-Patent Document 1, by implementing the behavior value function with a convolutional neural network, for example, when an object in an image rotates, the value of the behavior value function also changes according to the rotation of the object.
[0005] On the other hand, for example, there may be a case where a robot executes a task such as moving an object. In this case, for example, when a robot executes reinforcement learning, the robot needs to learn a movement path of the object. For example, when a robot executes a peg-in-hole task of inserting a peg, which is an example of an object, into a hole, the robot needs to learn what movement path to move the peg to the hole.
[0006] In Non-Patent Document 1, a behavior value function considering the rotational symmetry of the object itself is realized by a convolutional neural network. Therefore, even if the technique of Non-Patent Document 1 is used, only a learned model considering the rotational symmetry of the object itself can be obtained, and the learned model is a learned model used when gripping an object. Even if an attempt is made to obtain a learned model used when executing a task of moving an object using the technique of Non-Patent Document 1, a huge computational cost is required, and thus the learned model cannot be efficiently generated.
[0007] Note that in order to generate a learned model used when executing a task of moving an object, it is necessary to use a huge amount of learning data representing a movement history when the object is actually moved. However, when generating a learned model using such a huge amount of learning data, there is a problem that a huge computational cost is required and the learned model cannot be efficiently generated.
[0008] The present disclosure has been made in view of the above points, and an object thereof is to efficiently generate a learned model used when executing a task of moving an object from a starting point to an ending point.
Means for Solving the Problem
[0009] In order to achieve the above object, a learned model generation device according to the present disclosure is a learned model generation device that generates a learned model for controlling the movement of an object, and includes a first data representing a movement history of a learning object from a first starting point to an ending point, and a learning acquisition unit that acquires a pair with second data representing a movement history of a learning object from a second starting point to an ending point, and a second loss function representing a difference between an output value when the first data is input to a learning model and an output value when the second data is input to the learning model is added to a first loss function used when executing reinforcement learning, and a setting unit that sets an overall loss function for generating the learned model, and based on the pair acquired by the learning acquisition unit, a learning unit that generates the learned model that outputs movement data representing a displacement of the position of the object when the position data of the object is input by performing reinforcement learning on the learning model so that the output value of the overall loss function becomes smaller. A learned model generation device comprising:
[0010] Further, the learned model generation method of the present disclosure is a learned model generation method for generating a learned model for controlling the movement of an object, which obtains a pair of first data representing the movement history of a learning object from a first starting point to an end point and second data representing the movement history of the learning object from a second starting point to the end point, and adds a second loss function representing the difference between the output value when the first data is input to the learning model and the output value when the second data is input to the learning model to a first loss function used when performing reinforcement learning, thereby setting an overall loss function for generating the learned model, and based on the obtained pair, the learning model is subjected to reinforcement learning so that the output value of the overall loss function becomes smaller, thereby generating the learned model that outputs movement data representing the displacement of the position of the object when the position data of the object is input, and is a learned model generation method executed by a computer for the process.
[0011] Further, the learned model generation program of the present disclosure is a learned model generation program for generating a learned model for controlling the movement of an object, which obtains a pair of first data representing the movement history of a learning object from a first starting point to an end point and second data representing the movement history of the learning object from a second starting point to the end point, and adds a second loss function representing the difference between the output value when the first data is input to the learning model and the output value when the second data is input to the learning model to a first loss function used when performing reinforcement learning, thereby setting an overall loss function for generating the learned model, and based on the obtained pair, the learning model is subjected to reinforcement learning so that the output value of the overall loss function becomes smaller, thereby generating the learned model that outputs movement data representing the displacement of the position of the object when the position data of the object is input, and is a learned model generation program for causing a computer to execute the process.
Advantages of the Invention
[0012] According to the learned model generation device, control device, learned model generation method, and learned model generation program of the present disclosure, a learned model used when executing a task of moving an object from a starting point to an ending point can be efficiently generated.
Brief Description of the Drawings
[0013]
Figure 1
Figure 2
Figure 3
Figure 4
Figure 5
Figure 6
Figure 7
Figure 8
Figure 9
Figure 10
Figure 11
Figure 12
Figure 13
Modes for Carrying Out the Invention
[0014] Hereinafter, an example of an embodiment of the present disclosure will be described with reference to the drawings. In this embodiment, a control system equipped with the control device according to the present disclosure will be described as an example. In each drawing, the same or equivalent components and parts are given the same reference numerals. Also, the dimensions and ratios in the drawings are exaggerated for the convenience of explanation and may be different from the actual ratios.
[0015] FIG. 1 is a diagram for explaining this embodiment. As shown in FIG. 1, the control system 10 of this embodiment includes a robot 12 and a control device 14. The robot 12 includes an arm 16. The robot 12 executes a peg-in-hole task, which is a task of inserting a peg T into a hole H, by operating the arm 16 according to a control signal output from the control device 14.
[0016] FIG. 2 is a diagram for explaining the tip portion of the arm 16. As shown in FIG. 2, the tip portion of the arm 16 includes an arm body 16A and a gripper 16B. Note that the arm body 16A and the gripper 16B are connected by a spring 16C, which is an example of an elastic body. The arm body 16A and the gripper 16B normally operate integrally. And, for example, when the peg T contacts the floor surface F of the hole H, the arm body 16A and the gripper 16B are controlled to be connected by the spring 16C. Thereby, for example, when the peg T contacts the floor surface F, the arm body 16A and the gripper 16B are connected by the spring 16C, making it easier to search for the position of the hole H.
[0017] Note that when detecting whether or not the peg T has contacted the floor surface F, for example, a force sensor (not shown) described later is installed on the arm 16. When a force is detected by the force sensor due to the peg T contacting the floor surface F, the arm body 16A and the gripper 16B are controlled to be connected by the spring 16C.
[0018] Figures 3 and 4 are diagrams for explaining the trajectories when moving the peg T toward the hole H. As shown in FIG. 3, the trajectory m1 and the trajectory m2 when moving the peg T toward the hole H are in a line-symmetric relationship. Specifically, the trajectory obtained by inverting the trajectory m1 in FIG. 3 with respect to the Y-axis corresponds to the trajectory m2 in FIG. 3. Further, as shown in FIG. 4, the trajectories m3, m4, m5, and m6 are in a point-symmetric relationship centered on the hole H.
[0019] When teaching the robot 12 to perform the operation of the peg-in-hole task as shown in FIGS. 3 and 4, for example, it is possible to use a known reinforcement learning algorithm. When generating a learned model for controlling the movement of the arm 16 of the robot 12 using reinforcement learning, a large amount of learning data is required.
[0020] Therefore, in the present embodiment, learning data is increased by converting the actually obtained data. For example, by inverting the trajectory m1 of the peg T shown in FIG. 3 with respect to the Y-axis, the trajectory m2 is generated. Then, the movement history corresponding to the trajectory m1 and the movement history corresponding to the trajectory m2 are set as learning data. Further, for example, by performing a rotation centered on the hole H on the trajectory m3 of the peg T shown in FIG. 4, the trajectories m4, m5, and m6 are generated. Then, the movement history corresponding to each generated trajectory is set as learning data.
[0021] However, when generating a learned model using the learning data increased as described above, since each of the increased large amount of learning data is different and separate data, there is a problem that the calculation cost increases and the learned model cannot be efficiently generated.
[0022] Therefore, in this embodiment, by performing reinforcement learning on the assumption that the movement history before conversion and the movement history after conversion among the plurality of increased learning data are the same, a learned model used for the operation of the robot 12 to move the object is generated. Specifically, a new auxiliary loss function is added to the existing loss function used in reinforcement learning. This auxiliary loss function is a loss function that includes the difference between the output value when data representing the movement history before conversion is input and the output value when data representing the movement history after conversion is input. By generating a learned model so that the output value of this auxiliary loss function becomes small, it becomes possible to efficiently generate a learned model used when executing a task of moving an object from a starting point to an ending point.
[0023] The following will be specifically described.
[0024] <Framework of this embodiment> Conventionally, by using a known motion capture system, the relative positional relationship among the arm, hole, and peg of a robot has been sequentially detected, and the operation of the robot's arm has been controlled. Even when causing the robot to execute a peg-in-hole task, the operation of the robot's arm has been controlled using a motion capture system. When using a motion capture system, it is possible to directly observe the relative positional relationship among the arm, hole, and peg.
[0025] In contrast, in this embodiment, a force sensor (not shown) is installed in the arm 16 portion of the robot 12, and the operation of the arm 16 of the robot 12 is controlled based on the force sense value output from the force sensor. The value output from the force sensor is a three-axis force value and a three-axis moment value.
[0026] Therefore, in this embodiment, the framework of the Partially observable Markov decision process (POMDP) is adopted. The POMDP framework is a framework for cases where only a part of the state of the target system can be observed. In this embodiment, in order to control the operation of the arm 16 of the robot 12 based on the force sense value detected by the force sense sensor without directly observing the relative positional relationship between the arm 16, the hole H, and the peg T, this POMDP framework is adopted to control the operation of the arm 16 of the robot 12.
[0027] In addition, in this embodiment, the Deep reinforcement learning (DRL) approach and the Soft-Actor Critic (SAC) framework are adopted. SAC is disclosed, for example, in the following Reference 1.
[0028] Reference 1: T. Haarnoja, A. Zhou, P. Abbeel, and S. Levine, “Soft actor-critic: Off policy maximum entropy deep reinforcement learning with a stochastic actor,” in International Conference on Machine Learning, 2018, pp.1861-1870.
[0029] (Partially observable Markov decision process) The POMDP framework is defined by a tuple (S, A, Ω, T, R, O). S is the state space, A is the action space, and Ω is the observation space. T represents the dynamic function, R represents the reward value, and O is the observation function. After an action a is taken by an agent, state s changes to state s' according to the dynamic function T(s, a, s'). When transitioning from state s to state s', an observation o determined by the observation function O(a, s', o) is output. In order for an agent to behave optimally, a movement history h, which is time series data of observation o and action a, as shown in the following equation, is required. t It is necessary to determine appropriately.
[0030]
number
[0031] The movement history h is calculated so that the total reward J is maximized as follows: t Measure π(h t By generating R(s ), the agent can take appropriate action a according to the observation o. t ,a t ) is the state s of the agent at time t. t Action a t is the reward value obtained when taking . γ is the discount rate. E represents the expected value.
[0032]
number
[0033] In this embodiment, the movement history h t Measure π(h t ) and movement history h t The action value function Q(s t ,a t ) is learned. Details will be explained later.
[0034] (formulation) FIG. 5 is a diagram for explaining the augmentation of learning data. When X-Y coordinates with the hole H as the origin are set as shown in FIG. 5, assume that the coordinate data (x, y) of the starting point p0 of the peg T is the learning data actually obtained. In this case, let the function that performs the transformation with the X-axis as the axis of inversion be F x and the function that performs the transformation with the Y-axis as the axis of inversion be F y Also, let the function that performs a counterclockwise rotation about the origin be R n
[0035] Here, a transformation set G is defined with the transformation element g representing the transformation of the learning data as a constituent element. Also, an element e where the input is not transformed is defined. For example, when the transformation element g is F x , F x = {e, s x}, and s x represents the inversion operation with the X-axis as the axis of inversion. When the transformation element g is F x , as shown in FIG. 5, F x * p0 = {e * p0, s x * p0} = {p0, p1}. Similarly, when the transformation element g is F y , F x = {e, s y}, and s y represents the inversion operation with the Y-axis as the axis of inversion. As shown in FIG. 5, F y * p0 = {e * p0, s y * p0} = {p0, p5}. Here, for simplicity of notation, F x * F y is denoted as F xy . When the transformation element g is F xy , F xy * p0 = {p0, p1, p4, p5}.
[0036] Also, in this embodiment, a function R n that generates n planar rotations {0, 2π / n, 4π / n, ···, 2π(n - 1) / n} is defined. Note that e in this case corresponds to a rotation of 0 degrees. When the transformation element g is R n , R n By applying it to the coordinate data p0, data obtained by rotating p0 counterclockwise around the origin is generated. For example, R 4 *p0 = {p0, p2, p4, p6}. Furthermore, F xy and R 4 when combined, F xy *R 4 *p0 = {p0, p1, p2, p3, p4, p5, p6, p7}.
[0037] As described above, when learning data is generated by combining these functions, the following occurs. Also, the positional relationship of each learning data is the positional relationship as shown in FIG. 5.
[0038] F x *p0 = {p0, p1} F y *p0 = {p0, p5} R 4 *p0 = {p0, p2, p4, p6} F x *F y *p0 = F xy *p0 = {p0, p1, p4, p5} F xy *R 4 *p0 = {p0, p1, p2, p3, p4, p5, p6, p7}
[0039] Note that in this embodiment, |G| is defined as the size of the conversion set G. For example, |R n | = n, |F x | = |F y | = 2, |F xy | = 4, |F xy *R 4 | = 8.
[0040] (Variables set in this embodiment) In this embodiment, the observation o ∈ Ω is the three-dimensional position (t x , t y , t z ) of the arm main body 16A and the three-axis torque (τ x , τ y , τ z) and the three-axis forces (f x , f y , f z ), etc., are data. Also, the continuous action a = (δ x , δ y , δ z ) corresponds to the three-dimensional displacement of the arm main body 16A. In order to utilize the flexibility of the arm 16 by the spring 16C, it is preferable to adjust the position of the arm 16 of the robot 12 after bringing the peg T into contact with the floor surface F. The state s ∈ S is data including the relative posture between the peg T and the hole H, and s = p p2h is expressed as such. In this embodiment, since the state s is not directly observed, a partially observable Markov decision process is adopted. The observation o, the action a, the state s, and the reward r are represented by the following formula (1). As shown in formula (1), the reward r becomes 1 when the task is successful.
[0041]
Equation
[0042] (Proposed method) Figure 6 is a diagram for explaining the expansion of learning data considering symmetry. In Figure 6, the movement history of the arm 16 projected onto the XY plane is shown. The movement history h T 0 shown in Figure 6 is the actual movement history of the arm main body 16A of the robot 12 and is the movement history that successfully inserted the peg T into the hole H. Here, the transformation set G = F xy is selected, and when F T 0 acts on the movement history h xy , three movement histories (h ’ T 1 , h ’ T 2 , h ’ T 3 ) are generated. The movement history h T 0 and the three movement histories (h ’ T1 , h ’ T 2 , h ’ T 3 ) and the relationship between them is a line-symmetric relationship. These three movement histories (h ’ T 1 , h ’ T 2 , h ’ T 3 ) are used as learning data considering symmetry. Note that the movement history h T 0 corresponds to the movement history of the arm 16 of the robot 12 when the task is successful. The movement history when the task fails can also be used as learning data in the same way.
[0043] Also, FIG. 6 shows the symmetry of the action value function Q. When the movement history h t is given as the current movement history h T 0 , Q(h', t , a' t ) and Q(h', t 1 , a' t 1 ), Q(h', t 2 , a' t 2 ), and Q(h', t 3 , a' t 3 ) must be equal. Here, a' t i (i = 1, 2, 3) is a converted from a by the conversion set G t . Therefore, as shown in the following equation (2), it is necessary to generate an action value function Q that is invariant to the conversion by the conversion set G.
[0044]
Equation
[0045] Also, the action a at time t+1 t+1 =π(h t ) where, when the movement history h t converted from is input to the policy π, it is possible to derive the following equation (3) in the same manner as the conversion of the action a t+1 .
[0046] [Equation]
[0047] Therefore, the conversion set G = F xy and the action at time t+1 corresponding to the given movement history (h ’ T 1 , h ’ T 2 , h ’ T 3 ) is (h ’ t+1 1 , h ’ t+1 2 , h ’ t+1 3 ). The actions a t+1 , a ’ t+1 1 , a ’ t+1 2 , and a ’ t+1 3 shown in FIG. 6 are represented by arrows, and it is shown that the arm 16 of the robot 12 is displaced in the direction of the arrow.
[0048] The policy π that outputs the action data a before conversion shown in FIG. 6 t+1 and the policy π that outputs the action data a ’ t+1 1 after conversion need to be equal. Similarly, the policy π that outputs the action data a t+1 before conversion and the policy π that outputs the action data a ’ t+1 2The policy π that outputs must be equal, and the action data a before conversion t+1 The policy π that outputs and the action data a after conversion ’ t+1 3 The policy π that outputs must be equal.
[0049] (Data augmentation considering symmetry) When the transformation element g ∈ G is given, a data augmentation formula as shown in the following formula (4) is defined.
[0050] [Number]
[0051] When the movement history h is transformed by the transformation element g t the transformation element g is applied to all observations o and all actions a within the movement history h. Here, the following definitions are made. t
[0052] (Definition 1) g*o: Similar to p0 shown in Figure 5, F xy is the Z-axis element (t z , τ z , f z ) does not change, and acts only on (t x , t y ), (τ x , τ y ), (f x , f y ). (Definition 2) g*a: Similar to g*o, F xy is the Z-axis element and does not change a z . Also, similar to p0 shown in Figure 5, F xy acts only on (t x , t y ), (τ x , τ y , (f x , f y ).
[0053] (Loss function considering transformation) In this embodiment, the action value function Q and the policy π are learned so that the above formulas (2) and (3) are satisfied. As described above, in this embodiment, the SAC disclosed in the above Reference 1 is used to learn the action value function Q and the policy π.
[0054] FIG. 7 is a diagram showing the recurrent SAC model in this embodiment. The recurrent SAC model is disclosed in the following Reference 2.
[0055] Reference 2: T. Ni, B. Eysenbach, and R. Salakhutdinov, “Recurrent model-free RL can be a strong baseline for many POMDPs,” in International Conference on Machine Learning, 2022, pp. 16691-16723.
[0056] "ACTOR" in FIG. 7 is a learning model for learning the policy π (hereinafter, also simply referred to as "actor"), and "CRITICS" in FIG. 7 is a learning model for learning the action value function Q (hereinafter, also simply referred to as "critic"). "RNN" in FIG. 7 represents a recurrent neural network, "MLP" represents a multi-layer perceptron, and "Embedder" is a neural network model for embedding the observation o and the action a.
[0057] There are two differences between the recurrent SAC model of this embodiment and the SAC model disclosed in the above Reference 2. The first difference is that the learning data is increased by performing data augmentation as described above. The second difference is to introduce an auxiliary loss function such that the above formulas (2) and (3) are satisfied. Hereinafter, the auxiliary function will be described. Note that the following auxiliary loss functions L sym C (Q) and the auxiliary loss function L sym a (π) are an example of the second loss function of the present disclosure.
[0058] (Auxiliary loss function for the critic) As shown in the above formula (2), in this embodiment, it is necessary to train the critic so that the output value of the action value function Q for (h, a) before conversion is equal to the output value of the action value function Q for (g*h, g*a) converted by g included in the conversion set G. Therefore, in this embodiment, the critic is trained so that the difference between the action value function Q(h, a) before conversion and the action value function Q(g*h, g*a) after conversion becomes small. Specifically, the auxiliary loss function L sym C (Q) is minimized to train the critic.
[0059] [Number]
[0060] Note that, as shown in FIG. 7, when the model disclosed in the above Reference 2 is adopted, the output values of two types of action value functions Q1 and the output value of the action value function Q2 are generated. When the SAC model of Reference 2 is not adopted, one value of the action value function Q is output from the critic.
[0061] (Auxiliary loss function for the actor) The actor in SAC outputs the mean μ a of the action a and the standard deviation σ a of the action a. Therefore, in this embodiment, the actor is trained so that the difference between the mean μ h a of the action a before conversion and the mean μ g*h a of the action a after conversion becomes small. Specifically, the auxiliary loss function L sym a (π) is minimized to train the actor.
[0062] [Number]
[0063] Note that regarding the average μ of action a and the standard deviation σ of action a output by the actor in SAC, the details are disclosed in Reference 3 below. a and the standard deviation σ of action a a Regarding the point of outputting them, the details are disclosed in Reference 3 below.
[0064] Reference 3: H. Nguyen, A. Baisero, D. Wang, C. Amato, and R. Platt, “Leveraging fully observable policies for learning under partial observability,” in Conference on Robot Learning, 2022.
[0065] (Control System 10) FIG. 8 is a block diagram showing a schematic configuration of the control system 10 of the present embodiment. As shown in FIG. 8, the control system 10 includes a force sensor 11, a robot 12, and a control device 14. The control device 14 generates a learned model for controlling the operation of the arm 16 of the robot 12. Further, the control device 14 controls the operation of the arm 16 of the robot 12 using the generated learned model.
[0066] The force sensor 11 is attached to the arm 16 of the robot 12 and sequentially detects the three-axis torque (τ x , τ y , τ z ) and the three-axis force (f x , f y , f z ) which are force sense values. Then, the force sensor 11 outputs the obtained force sense values to the control device 14.
[0067] Note that among the observations o shown in the above formula (1), the three-dimensional position (t x , t y , t z) can be easily calculated from the displacement represented by the action a. Therefore, the force sensing values (τ x , τ y , τ z , f x , f y , f z ) obtained by the force sensor 11 and the three-dimensional position (t x , t y , t z ) of the arm body 16A are used to set the observation o shown in the above formula (1).
[0068] The robot 12 is a robot as shown in FIG. 1, and executes a peg-in-hole task of inserting the peg T into the hole H by operating the arm 16.
[0069] FIG. 9 is a block diagram showing the hardware configuration of the control device 14 according to the present embodiment. As shown in FIG. 9, the control device 14 includes a CPU (Central Processing Unit) 42, a memory 44, a storage device 46, an input / output I / F (Interface) 48, a storage medium reader 50, and a communication I / F 52. Each component is connected to be communicable with each other via a bus 54.
[0070] The storage device 46 stores a learned model generation program and a control program for executing each process described later. The CPU 42 is a central processing unit that executes various programs and controls each component. That is, the CPU 42 reads a program from the storage device 46 and executes the program using the memory 44 as a work area. The CPU 42 performs control of the above components and various arithmetic processes according to the program stored in the storage device 46.
[0071] The memory 44 is composed of a RAM (Random Access Memory) and temporarily stores programs and data as a working area. The storage device 46 is composed of a ROM (Read Only Memory), an HDD (Hard Disk Drive), an SSD (Solid State Drive), etc., and stores various programs including an operating system and various data.
[0072] The input / output I / F 48 is an interface for inputting data from an external device and outputting data to an external device. Also, for example, input devices for performing various inputs such as a keyboard and a mouse, and output devices for outputting various information such as a display and a printer may be connected. By adopting a touch panel display as the output device, it may function as an input device.
[0073] The storage medium reader 50 reads data stored in various storage media such as a CD (Compact Disc)-ROM, a DVD (Digital Versatile Disc)-ROM, a Blu-ray disc, a USB (Universal Serial Bus) memory, etc., and writes data to the storage medium.
[0074] The communication I / F 52 is an interface for communicating with other devices, and for example, standards such as Ethernet (registered trademark), FDDI, Wi-Fi (registered trademark), etc. are used.
[0075] Next, the functional configuration of the control device 14 will be described. As shown in FIG. 8, the control device 14 functionally includes a learning acquisition unit 17, a setting unit 18A, a learning unit 18B, an acquisition unit 20, and a control unit 22. Also, a data storage unit 24 and a learned model storage unit 26 are provided in a predetermined storage area of the control device 14. Each functional configuration is realized by the CPU 42 reading each program stored in the storage device 46, expanding it in the memory 44, and executing it.
[0076] The force sense value detected by the force sense sensor 11 is stored in the data storage unit 24. Further, control data when the arm 16 of the robot 12 operates is stored in the data storage unit 24. For example, the action (δ x , δ y , δ z ) when the arm 16 operates is stored.
[0077] The learned model storage unit 26 stores the learned policy π and the learned action value function Q generated by the process described later.
[0078] First, the learning acquisition unit 17 and the learning unit 18B generate a learned policy π which is a learned model for controlling the operation of the arm 16 of the robot 12.
[0079] Various data including the movement history h t when the arm 16 of the robot 12 is executing the peg-in-hole task is stored in the data storage unit 24. The movement history data at this time includes data when the peg-in-hole task is successful and data when the peg-in-hole task fails. Note that the data representing the movement history obtained by the actual operation of the arm 16 of the robot 12 is hereinafter simply referred to as "first data". The first data is the movement history when the arm 16, which is an example of the object, is moved from the first starting point p0 to the hole H which is the end point.
[0080] The learning acquisition unit 17 reads a plurality of first data from the data storage unit 24. Then, for each of the plurality of first data, the learning acquisition unit 17 generates each of the second data representing the movement history having symmetry with the first data by applying the transformation element g included in the above-described transformation set G. As shown in FIGS. 5 and 6, the second data is the movement history when the arm main body 16A, which is an example of the object, is moved from the second starting points p1, p2, p3, p4, p5, p6, p7 to the hole H which is the end point. In this way, the learning acquisition unit 17 acquires a plurality of pairs of the first data and the second data. As a result, the learning data increases, and it becomes possible to generate a highly accurate learned model.
[0081] The setting unit 18A adds the auxiliary loss function L A (Q) shown in the above formula (5) to the known loss function L sym C used when executing reinforcement learning, and sets the overall loss function L A C for generating a learned model. As shown in the above formula (5), the auxiliary loss function L sym C (Q) is a loss function representing the difference between the output value when the first data is input to the critic and the output value when the second data is input to the critic. For example, the sum of the loss function L A and the auxiliary loss function L sym C (Q) is set as the overall loss function L A C .
[0082] Also, the setting unit 18A adds the auxiliary loss function L B (π) shown in the above formula (6) to the known loss function L sym a used when executing reinforcement learning, and sets the overall loss function L B a for generating a learned model. As shown in the above formula (6), the auxiliary loss function L sym a(π) is a loss function representing the difference between the output value when the first data is input to the actor and the output value when the second data is input to the actor. For example, the loss function L B and the auxiliary loss function L sym a (π) are set as the overall loss function L B a .
[0083] Note that the known loss functions L A and L B are, for example, loss functions used in the SAC algorithm. The known loss functions L A and L B are an example of the first loss function of the present disclosure.
[0084] The learning unit 18B learns the above-described actor and critic by performing reinforcement learning based on a plurality of pairs acquired by the learning acquisition unit 17. Note that when performing reinforcement learning, the learning unit 18B performs reinforcement learning on the learning model assuming that the first data and the second data in the pair are the same. Thereby, a learned policy π corresponding to the actor and an action value function Q corresponding to the critic are generated.
[0085] Note that the learned policy π is a learned model that outputs movement data representing the displacement (δ x , δ y , δ z ) of the position of the arm 16 when an observation o including the three-dimensional position (t x , t y , t z ) of the arm 16 is input. The learned policy π is used when operating the arm 16 of the robot 12.
[0086] The learning unit 18B performs reinforcement learning so that the overall loss function L A C set by the setting unit 18A becomes small. Also, when performing reinforcement learning, the learning unit 18B performs reinforcement learning so that the overall loss function L B aBy performing reinforcement learning so as to minimize [the relevant value], a learned policy is generated. As described above, the policy π corresponds to the actor, and the action-value function Q corresponds to the critic.
[0087] Then, the learning unit 18B stores the learned policy π corresponding to the actor and the learned action-value function Q corresponding to the critic in the learned model storage unit 26.
[0088] When the learned policy π is stored in the learned model storage unit 26, it becomes possible to control the operation of the arm 16 of the robot 12 using the learned policy π. For this reason, the acquisition unit 20 and the control unit 22 control the operation of the arm 16 of the robot 12 using the learned policy π stored in the learned model storage unit 26.
[0089] The acquisition unit 20 acquires an observation o including the three-dimensional position (t x , t y , t z ) of the arm 16.
[0090] The control unit 22 reads out the learned policy π stored in the learned model storage unit 26. Then, the control unit 22 inputs the observation o including the three-dimensional position (t x , t y , t z ) of the arm 16 acquired by the acquisition unit 20 into the learned policy π. From the learned policy π, movement data representing the displacement (δ x , δ y , δ z ) of the position of the arm 16 is output. The control unit 22 acquires the movement data output from the learned policy π and controls the position of the arm 16 based on the movement data. Specifically, the control unit 22 outputs a control signal to the arm 16 of the robot 12 such that the displacement (δ x , δ y , δ z ) of the position of the arm 16 is realized.
[0091] Next, the operation of the control system 10 according to the present embodiment will be described.
[0092] First, the movement history of the arm 16 of the robot 12 while the arm 16 is performing the peg-in-hole task is collected and input to the control device 14. Then, the data regarding the movement history is stored in the data storage unit 24. When the control device 14 receives a predetermined instruction signal, the CPU 42 of the control device 14 reads out the learned model generation program from the storage device 46, expands it in the memory 44, and executes it. As a result, the CPU 42 functions as each functional configuration of the control device 14, and the learned model generation process shown in FIG. 10 is executed.
[0093] In step S100, the learning acquisition unit 17 acquires a plurality of first data by reading out the plurality of first data stored in the data storage unit 24.
[0094] In step S102, the setting unit 18A generates each of the second data representing the movement history having symmetry with the first data by applying the conversion element g included in the above-described conversion set G to each of the plurality of first data acquired in step S100. In this way, the setting unit 18A acquires a plurality of pairs of the first data and the second data.
[0095] In step S103, the setting unit 18A adds the auxiliary loss function L A shown in the above formula (5) to the known loss function L sym C (Q) used when performing reinforcement learning to set the overall loss function L A C for generating the learned model. Also, in step S103, the setting unit 18A adds the auxiliary loss function L B shown in the above formula (6) to the known loss function L sym a (π) used when performing reinforcement learning to set the overall loss function L B a for generating the learned model.
[0096] In step S104, the learning unit 18B learns a critic which is an action value function Q and an actor which is a policy π by executing reinforcement learning. Specifically, the learning unit 18B learns the critic which is an action value function so that the overall loss function L A C becomes smaller according to a known SAC algorithm. Also, in step S104, the learning unit 18B learns the actor which is a policy so that the overall loss function L B a becomes smaller according to a known SAC algorithm.
[0097] In step S106, the learning unit 18B stores the learned action value function Q and the learned policy π obtained in step S104 in the learned model storage unit 26.
[0098] Next, when the control device 14 receives a predetermined instruction signal, the control device 14 executes the control process shown in FIG. 11.
[0099] In step S200, the acquisition unit 20 acquires an observation o at time t including the three-dimensional position (t x , t y , t z ) of the arm 16, and an action a at time t - 1 t =(t x , t y , t z , τ x , τ y , τ z , f x , f y , f z ). t-1 =(δ x , δ y , δ z ).
[0100] In step S202, the control unit 22 reads out the learned policy π stored in the learned model storage unit 26. Then, in step S202, the control unit 22 uses the observation o at time t and the action a at time t - 1 acquired in step S200 t and the action a at time t - 1 t-1and input it into the learned policy π. From the learned policy π, movement data representing the displacement of the position of the arm 16 (δ x , δ y , δ z ) is output. Therefore, the control unit 22 acquires the movement data output from the learned policy π.
[0101] In step S204, the control unit 22 controls the position of the arm 16 based on the movement data acquired in step S202. Specifically, the control unit 22 outputs a control signal to the arm 16 of the robot 12 such that the displacement of the position of the arm 16 (δ x , δ y , δ z ) is realized.
[0102] As described above, the control device according to the present embodiment generates a learned model for controlling the movement of an object. The control device acquires a pair of first data representing the movement history of the learning object from the first starting point to the end point and second data representing the movement history of the learning object from the second starting point to the end point. Further, the control device adds an auxiliary loss function representing the difference between the output value when the first data is input to the learning model and the output value when the second data is input to the learning model to a known loss function used when performing reinforcement learning, thereby setting an overall loss function for generating the learned model. When performing reinforcement learning based on the acquired pair, the control device causes the learning model to perform reinforcement learning so that the output value of the overall loss function becomes small, thereby generating a learned model that outputs movement data representing the displacement of the position of the object when the position data of the object is input. Thereby, a learned model used when executing the task of moving the object from the starting point to the end point can be efficiently generated.
[0103] Specifically, the control device of the present embodiment generates a learned model for controlling the operation of the arm of a robot, which is an example of an object. The first data and the second data for controlling the operation of the robot arm are data representing the displacement of the arm position and the force sense value output from the force sense sensor installed on the arm. Further, the second data is data obtained by converting the first data and has symmetry with the first data. The control device generates a learned model that outputs the movement data of the arm when the position data of the arm is input.
[0104] According to the control device of the present embodiment, by executing reinforcement learning assuming that the first data representing the movement history of the arm from the first start point to the end point and the second data representing the movement history of the arm from the second start point to the end point are the same, the calculation cost for generating the learned model is reduced, so that the learned model can be efficiently generated.
[0105] Further, according to the control device of the present embodiment, by generating the second data by converting the first data and using the first data and the second data as learning data, it is possible to generate a learned model with higher accuracy.
[0106] Further, according to the control device of the present embodiment, the observation o=(t x ,t y ,t z ,τ x ,τ y ,τ z ,f x ,f y ,f z ), and information regarding the relative posture between the peg T and the hole H becomes unnecessary. Therefore, if the force sense value obtained from the force sense sensor 11 and the position of the arm main body 16A can be acquired, it is possible to execute a task of moving an object such as a peg-in-hole task without using motion capture technology or the like.
Example
[0107] Next, the embodiments will be described. In this embodiment, a simulation is performed to verify the effectiveness of the proposed method. In this simulation, simulations of peg-in-hole are executed when the shape of the hole is triangular, square, pentagonal, hexagonal, and circular.
[0108] Figure 12 is a diagram showing the simulation results of this embodiment. "SAC-State" shown in Figure 12 is the result when the policy is learned based on state s. Here, state s is information including the relative positional relationship between the peg and the hole. Also, "SAC-Obs" is the result when the policy is learned only with observation o. "RSAC" corresponds to "SAC-Obs" when the recurrent SAC model is used. These three agents, "SAC-State", "SAC-Obs", and "RSAC", are learned without using data symmetry and data augmentation.
[0109] The proposed method of this embodiment corresponds to "RSAC-Aug-Aux". Note that "RSAC-Equi" corresponds to the method disclosed in the following Reference 4. The method of the following Reference 4 is a POMDP-based method.
[0110] Reference 4: H. Nguyen, A. Baisero, D. Klee, D. Wang, R. Platt, and C. Amato, “Equivariant reinforcement learning under partial observability,” in Conference on Robot Learning, 2023. [Online]. Available: https: / / openreview.net / forum?id=AnDDMQgM7-
[0111] The horizontal axis "Environment Step" in Figure 12 represents the number of learning times, and "Evaluation Success Rate" is the success rate for 6 attempts of the peg-in-hole. As shown in Figure 12, it can be seen that for any shape of the hole, the proposed method of this embodiment has the highest success rate for "RSAC-Aug-Aux".
[0112] Also, Figure 13 is a comparison of the results when both data symmetry and data augmentation are performed, and the results when only one of them is performed. "Aug-Aux" represents the result when both are performed, "Aug" represents the result when only data augmentation is performed, and "Aux" represents the result when only data symmetry is utilized.
[0113] As can be seen from Figures 12 and 13, it can be understood that by using the method of this embodiment, it is possible to accurately and efficiently generate the learned policy used when moving the object.
[0114] In addition, in the above embodiment, the case where the object for learning and the object are the peg T or the arm main body 16A is described as an example, but it is not limited thereto. Any object that can be moved may be used. For example, the method of this embodiment is applicable not only to the peg-in-hole task, but also to picking tasks or navigation tasks, etc. Also, the method of this embodiment is applicable to tasks that move the object by autonomous driving. Further, this embodiment may be applied to the control of the operation of other parts different from the robot arm, the autonomous driving of a mobile robot, autonomous flight, or the control of autonomous navigation, etc., for other movements.
[0115] In addition, in the above embodiment, the case where data augmentation for generating the second data from the first data is performed has been described as an example, but the present invention is not limited thereto. It is not necessary to perform data augmentation. For example, among a plurality of data, if it is appropriate to treat a certain data A and another data B in the same column in the action value function Q and the policy π, when executing reinforcement learning, a certain data A and another data B may be treated as being the same. Further, for example, the second data does not have to be data having symmetry with the first data. For example, the second data may be data obtained by performing some conversion process on the first data. For example, the second data may be data obtained by temporally converting the first data. Further, for example, in the case of a video, the second data may be data obtained by playing the video in reverse.
[0116] In addition, in the above embodiment, the case where the action value function Q and the policy π are learned according to the SAC algorithm has been described as an example, but the present invention is not limited thereto. The action value function Q and the policy π may be learned according to other reinforcement learning algorithms.
[0117] In addition, in the above embodiment, the movement history h was data including the action a and the state o, but the present invention is not limited thereto. The movement history h can be changed as appropriate. For example, only the position of the object may be included.
[0118] In addition, in the above embodiment, the case where the first data and the second data are the displacement of the position of the arm and the force sense value output from the force sensor installed on the arm has been described as an example, but the present invention is not limited thereto. For example, the first data and the second data may be data including either the displacement of the position of the arm or the force sense value output from the force sensor installed on the arm.
[0119] In the above-described embodiment, the case where the control device 14 executes both the learned model generation process of FIG. 10 and the control process of FIG. 11 has been described as an example, but the present invention is not limited thereto. For example, a learned model generation device realized by a computer different from the control device 14 may be prepared, and the learned model generation device may execute the learned model generation process of FIG. 10, and the control device 14 may execute the control process of FIG. 11. In this case, the learned model generation device includes at least the above-described learning acquisition unit 17, setting unit 18A, and learning unit 18B.
[0120] In the above-described embodiment, each process executed by the CPU by reading software (program) may be executed by various processors other than the CPU. Examples of the processor in this case include a PLD (Programmable Logic Device) whose circuit configuration can be changed after manufacturing, such as an FPGA (Field-Programmable Gate Array), and a dedicated electric circuit which is a processor having a circuit configuration designed specifically for executing specific processes, such as an ASIC (Application Specific Integrated Circuit). Further, each process may be executed by one of these various processors, or may be executed by a combination of two or more processors of the same type or different types (for example, a combination of a plurality of FPGAs, and a combination of a CPU and an FPGA, etc.). More specifically, the hardware structure of these various processors is an electric circuit combining circuit elements such as semiconductor elements.
[0121] In the above-described embodiment, the mode in which each program is stored (installed) in advance in the storage device has been described, but the present invention is not limited thereto. The program may be provided in a form stored in a storage medium such as a CD-ROM, DVD-ROM, Blu-ray Disc, USB memory, etc. Further, the program may be in a form downloaded from an external device via a network.
[0122] (Supplementary Note) Hereinafter, aspects of the present disclosure will be appended.
[0123] (Appendix 1) A learned model generation device that generates a learned model for controlling the movement of an object, a learning acquisition unit that acquires a pair of first data representing the movement history of a learning object from a first starting point to an end point and second data representing the movement history of the learning object from a second starting point to the end point, a setting unit that sets an overall loss function for generating the learned model by adding a second loss function representing the difference between the output value when the first data is input to the learning model and the output value when the second data is input to the learning model to a first loss function used when performing reinforcement learning, a learning unit that generates the learned model that outputs movement data representing the displacement of the position of the object when the position data of the object is input by performing reinforcement learning on the learning model so that the output value of the overall loss function becomes small based on the pair acquired by the learning acquisition unit, A learned model generation device comprising the above.
[0124] (Appendix 2) The second data is data obtained by converting the first data. The learned model generation device according to Appendix 1. The learning object is a robot arm, The first data and the second data are data representing the displacement of the position of the arm and the force sense value output from a force sense sensor installed on the arm, The second data is data obtained by converting the first data and has symmetry with the first data, The learning unit generates the learned model that outputs the movement data of the arm when the position data of the arm is input. The learned model generation device according to Appendix 1.
[0125] (Appendix 3) The learning object is a robot arm, The first data and the second data are data including at least one of displacement of the position of the arm and a force sense value output from a force sensor installed on the arm. The second data is data obtained by converting the first data and is data having symmetry with the first data. The learning unit generates the learned model that outputs the movement data of the arm when the position data of the arm is input. The learned model generation device according to Appendix 1 or Appendix 2.
[0126] (Appendix 4) The relationship between the movement history represented by the first data and the movement history represented by the second data is a line-symmetric or point-symmetric relationship. The learned model generation device according to any one of Appendices 1 to 3.
[0127] (Appendix 5) The second loss function is a loss function including a difference between an output value of the action value function when the first data is input to the action value function in reinforcement learning and an output value of the action value function when the second data is input to the action value function. The learned model generation device according to any one of Appendices 1 to 4.
[0128] (Appendix 6) The learning model and the learned model correspond to a policy in reinforcement learning. The second loss function is a loss function including a difference between an output value of the policy when the first data is input to the policy and an output value of the policy when the second data is input to the policy. The learned model generation device according to any one of Appendices 1 to 5.
[0129] (Appendix 7) The policy in reinforcement learning is an actor in the Soft Actor-Critic algorithm. The action value function in reinforcement learning is the critic in the Soft Actor-Critic algorithm, The second loss function includes the difference between the output value of the critic when the first data is input to the critic and the output value of the actor when the second data is input to the critic, and the difference between the output value of the actor when the first data is input to the actor and the output value of the actor when the second data is input to the actor, The learning unit, When performing reinforcement learning according to the Soft Actor-Critic algorithm, The action value function corresponding to the critic is learned and the policy corresponding to the actor is learned so that the overall loss function becomes smaller, and the learned model corresponding to the actor is generated. The learned model generation device according to any one of Appendices 1 to 6.
[0130] (Appendix 8) An acquisition unit that acquires position data of an object, By inputting the position data of the object acquired by the acquisition unit to the learned model generated by the learned model generation device according to any one of Appendices 1 to 7, the movement data of the object is acquired, and the position of the object is controlled based on the movement data of the object. A control unit, A control device comprising the same.
[0131] (Appendix 9) A learned model generation method for generating a learned model for controlling the movement of an object, By adding a second loss function representing the difference between the output value when the first data is input to the learning model and the output value when the second data is input to the learning model to the first loss function used when performing reinforcement learning, an overall loss function for generating the learned model is set, Based on the obtained pair, by performing reinforcement learning on the learning model so that the output value of the overall loss function becomes smaller, a learned model is generated that outputs movement data representing the displacement of the position of the object when the position data of the object is input. A method for generating a learned model for a computer to execute a process.
[0132] (Appendix 10) For the first loss function used when performing reinforcement learning, by adding a second loss function representing the difference between the output value when the first data is input to the learning model and the output value when the second data is input to the learning model, an overall loss function for generating the learned model is set. Based on the obtained pair, by performing reinforcement learning on the learning model so that the output value of the overall loss function becomes smaller, a learned model is generated that outputs movement data representing the displacement of the position of the object when the position data of the object is input. A learned model generation program for causing a computer to execute a process.
Explanation of Signs
[0133] 10 Control system 11 Force sensor 12 Robot 14 Control device 16 Arm 16A Arm body 16B Gripper 16C Spring 17 Learning acquisition unit 18A Setting unit 18B Learning unit 20 Acquisition unit 22 Control unit 24 Data storage unit 26 Learned model storage unit
Claims
1. A learned model generation device that generates a learned model for controlling the movement of an object, comprising: a learning acquisition unit that acquires a pair of first data representing the movement history of a learning object from a first starting point to an end point and second data representing the movement history of the learning object from a second starting point to the end point; a setting unit that sets an overall loss function for generating the learned model by adding a second loss function representing a difference between an output value when the first data is input to the learning model and an output value when the second data is input to the learning model, with respect to a first loss function used when performing reinforcement learning; a learning unit that generates the learned model that outputs movement data representing the displacement of the position of the object when the position data of the object is input, by performing reinforcement learning on the learning model so that the output value of the overall loss function becomes small, based on the pair acquired by the learning acquisition unit; A learned model generation device comprising the above.
2. The second data is data obtained by converting the first data. The learned model generation device according to Claim 1.
3. The learning object is a robot arm. The first data and the second data are data including at least one of the displacement of the position of the arm and the force sense value output from a force sense sensor installed on the arm. The second data is data obtained by converting the first data and has symmetry with the first data. The learning unit generates the learned model that outputs the movement data of the arm when the position data of the arm is input. The learned model generation device according to Claim 1 or Claim 2.
4. The relationship between the movement history represented by the first data and the movement history represented by the second data is a line-symmetric or point-symmetric relationship. The learned model generation device according to Claim 1 or Claim 2.
5. The second loss function is a loss function including a difference between the output value of the action value function when the first data is input to the action value function in reinforcement learning and the output value of the action value function when the second data is input to the action value function. The learned model generation device according to Claim 1 or Claim 2.
6. The learning model and the learned model correspond to policies in reinforcement learning. The second loss function is A loss function including the difference between the output value of the policy when the first data is input to the policy and the output value of the policy when the second data is input to the policy. The learned model generation device according to claim 1 or claim 2.
7. The policy in reinforcement learning is the actor in the Soft Actor-Critic algorithm, The action value function in reinforcement learning is the critic in the Soft Actor-Critic algorithm, The second loss function includes the difference between the output value of the critic when the first data is input to the critic and the output value of the actor when the second data is input to the critic, and the difference between the output value of the actor when the first data is input to the actor and the output value of the actor when the second data is input to the actor. The learning unit, When performing reinforcement learning according to the Soft Actor-Critic algorithm, The action value function corresponding to the critic is learned and the policy corresponding to the actor is learned so that the overall loss function becomes small, and the learned model corresponding to the actor is generated. The learned model generation device according to claim 1 or claim 2.
8. An acquisition unit that acquires position data of an object, By inputting the position data of the object acquired by the acquisition unit to the learned model generated by the learned model generation device according to claim 1 or claim 2, the movement data of the object is acquired, and the position of the object is controlled based on the movement data of the object. A control unit, A control device comprising:
9. A learned model generation method for generating a learned model for controlling the movement of an object, A pair of first data representing the movement history of a learning object from a first starting point to an end point and second data representing the movement history of a learning object from a second starting point to an end point is acquired, For the first loss function used when performing reinforcement learning, a second loss function representing the difference between the output value when the first data is input to the learning model and the output value when the second data is input to the learning model is added to set an overall loss function for generating the learned model. Based on the obtained pair, by performing reinforcement learning on the learning model so that the output value of the overall loss function becomes smaller, a learned model is generated that outputs movement data representing the displacement of the position of the object when the position data of the object is input. A method for generating a learned model in which a computer executes a process.
10. A learned model generation program for generating a learned model for controlling the movement of an object, Obtaining a pair of first data representing the movement history of a learning object from a first starting point to an end point and second data representing the movement history of a learning object from a second starting point to an end point, For a first loss function used when performing reinforcement learning, by adding a second loss function representing the difference between the output value when the first data is input to the learning model and the output value when the second data is input to the learning model, an overall loss function for generating the learned model is set. Based on the obtained pair, by performing reinforcement learning on the learning model so that the output value of the overall loss function becomes smaller, a learned model is generated that outputs movement data representing the displacement of the position of the object when the position data of the object is input. A learned model generation program for causing a computer to execute a process.