Learning systems and learning methods

JP2026131248APending Publication Date: 2026-08-14TOYOTA JIDOSHA KK
View PDF 1 Cites 0 Cited by

Patent Information

Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Filing Date
2025-02-03
Publication Date
2026-08-14

AI Technical Summary

Benefits of technology

【0007】 本開示によればマスタ装置を学習済みモデルによって動作させてスレーブ装置にタスクを実行させる場合に、学習済みモデルの再学習を効率的に行うことが可能な学習システム及び学習方法を提供できる。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2026131248000001_ABST
    Figure 2026131248000001_ABST
Patent Text Reader

Abstract

This system provides a learning system that enables efficient retraining of pre-trained models. [Solution] The first master device outputs a first observed value indicating the state of the first observed object using a trained model. The second master device has a second observed object that can be operated by an operator. The slave device has a third observed object, and the task is performed by the operation of the third observed object. The cooperative control unit controls the state of the first observed object, the state of the second observed object, and the state of the third observed object to correspond to each other. The learning unit uses command values ​​for the first master device and command values ​​for the slave device, calculated from the second observed value indicating the state of the second observed object and the third observed value indicating the state of the third observed object, to train the trained model so that the slave device can perform the task.
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] This disclosure relates to a learning system and a learning method. [Background technology]

[0002] A master-slave control system has been proposed. In relation to this technology, Patent Document 1 discloses a technique for operating a master actuator and a slave actuator by bilateral control and multilateral control, which is an extension of bilateral control. [Prior art documents] [Patent Documents]

[0003] [Patent Document 1] Japanese Patent Publication No. 2009-279699 [Overview of the project] [Problems that the invention aims to solve]

[0004] Patent Document 1 does not disclose how to perform training on a pre-trained model when operating a master device with a pre-trained model. Therefore, with the technology described in Patent Document 1, there is a risk that the pre-trained model cannot be efficiently retrained when operating a master device with a pre-trained model and having a slave device perform a task.

[0005] This disclosure provides a learning system and learning method that enable efficient retraining of a trained model when a master device is operated by a trained model and a slave device is made to perform a task. [Means for solving the problem]

[0006] The learning system according to the present disclosure includes a first master device that outputs a first observation value indicating the state of a first observation target by a learned model that has been learned in advance by machine learning, a second master device that has a second observation target that can operate by an operation by an operator, a slave device that has a third observation target and executes a task by operating the third observation target, a cooperative control unit that performs control so that the state of the first observation target of the first master device, the state of the second observation target of the second master device, and the state of the third observation target of the slave device correspond to each other, and a learning unit that uses a command value for the first master device and a command value for the slave device calculated from a second observation value indicating the state of the second observation target of the second master device and a third observation value indicating the state of the third observation target of the slave device to perform learning of the learned model so that the slave device executes the task.

Advantages of the Invention

[0007] According to the present disclosure, it is possible to provide a learning system and a learning method capable of efficiently performing relearning of a learned model when operating a master device by the learned model to cause a slave device to execute a task.

Brief Description of the Drawings

[0008] [Figure 1] It is a diagram for explaining an outline of the learning system according to Embodiment 1. [Figure 2] It is a diagram showing a configuration of the learning system according to Embodiment 1. [Figure 3] It is a diagram showing a configuration of the control device according to Embodiment 1. [Figure 4] It is a diagram illustrating a control block according to Embodiment 1. [Figure 5] It is a flowchart showing a learning method executed by the learning system according to Embodiment 1.

Modes for Carrying Out the Invention

[0009] This embodiment will be described below with reference to the drawings. However, the present invention is not limited to the following embodiments. Also, for clarity of explanation, the following description and drawings have been simplified as appropriate.

[0010] (Embodiment 1) Figure 1 is a diagram illustrating the outline of a learning system 1 according to Embodiment 1. The learning system 1 includes a first master device 10, a second master device 20, and a slave device 30. In the learning system 1, the first master device 10, the second master device 20, and the slave device 30 operate by multilateral control. That is, the learning system 1 is configured such that the operation of the first master device 10, the operation of the second master device 20, and the operation of the slave device 30 correspond to each other. In other words, the learning system 1 is configured by multilateral control such that the state of the first master device 10, the state of the second master device 20, and the state of the slave device 30 correspond to each other.

[0011] The second master device 20 and the slave device 30 have substantially similar hardware configurations. The second master device 20 and the slave device 30 function, for example, as robots. The second master device 20 has a robotic arm 22 (second object of observation). The robotic arm 22 has one or more joints. The slave device 30 has a robotic arm 32 (third object of observation). The robotic arm 32 has one or more joints.

[0012] Furthermore, the first master device 10 has a pre-trained model 12 that has been learned in advance by machine learning. The pre-trained model 12 simulates the state (movement) of the first observation target corresponding to the robot arms 22 and 32. In other words, the first observation target can be said to be a virtual robot arm. If the functions of the first master device 10 are incorporated into the functions of the robots of the second master device 20 and slave device 30, the state of the robot arm corresponding to the first observation target will correspond to the state of the robot arms 22 and 32.

[0013] The trained model 12 is pre-trained so that the slave device 30 can perform a predetermined task. For example, the trained model 12 is pre-generated by interactive imitation learning, in which the slave device 30 learns to perform a task by imitating the operations of the operator M. The first master device 10 outputs a first observation value indicating the state of the first observed object using the trained model 12. The "state of the first observed object" corresponds, for example, to the position and orientation of the robot arm corresponding to the first observed object. Specifically, the state of the first observed object may correspond to the joint angles of the robot arm corresponding to the first observed object. The state of the first observed object may also correspond to the end-effector position of the robot arm corresponding to the first observed object. Furthermore, the state of the first observed object may correspond to the orientation of the robot arm corresponding to the first observed object (rotation angle around the roll axis, rotation angle around the pitch axis, and rotation angle around the yaw axis). The same applies to the robot arm 22 (second observed object) and the robot arm 32 (third observed object). Furthermore, the state of robot arm 22 is observed, and a second observation value indicating the state of robot arm 22 is obtained. Similarly, the state of robot arm 32 is observed, and a third observation value indicating the state of robot arm 32 is obtained.

[0014] Through multilateral control, the state of the robot arm 32 of the slave device 30 follows the state indicated by the first observed value and the state of the robot arm 22 of the second master device 20 (the state indicated by the second observed value). Similarly, through multilateral control, the state of the robot arm 22 of the second master device 20 follows the state indicated by the first observed value and the state of the robot arm 32 of the slave device 30 (the state indicated by the third observed value). Further details will be described later.

[0015] The slave device 30 performs a task by operating the robot arm 32 (the third object of observation). The slave device 30 performs the task using the trained model 12 of the first master device 10. Specifically, the robot arm 32 of the slave device 30 moves the workpiece W so that it reaches a predetermined state. The task may be, for example, "to move the workpiece W to a predetermined position." In this case, the trained model 12 of the first master device 10 causes the robot arm 32 of the slave device 30 to move the workpiece W to the predetermined position. In this case, the trained model 12 is trained so that the robot arm 32 operates in such a way that the workpiece W moves to the predetermined position. Note that the above example of a task is merely illustrative. Therefore, the predetermined task is not limited to the above example.

[0016] Furthermore, operator M can operate the robot arm 22 of the second master device 20. Therefore, the second master device 20 is configured so that the robot arm 22 can be operated by operator M. When the robot arm 22 is operated by operator M, the second observed value described above may change. When the robot arm 22 is operated by operator M, the state of the robot arm 32 of the slave device 30 follows the state of the robot arm 22. For example, if the robot arm 32 of the slave device 30 is attempting to perform a task but is not performing the task well, operator M may operate the robot arm 22. This allows the robot arm 32 of the slave device 30 to move the workpiece W in accordance with the task. Therefore, the robot arm 32 of the slave device 30 can perform the task well.

[0017] In the example task described above, let's assume that the learned model 12 has already been learned, with the left position of the slave device 30 as the initial position A of the workpiece W. In this case, the robot arm 32 of the slave device 30 can operate according to the learned model 12 of the first master device 10 to move the workpiece W from the initial position A to a predetermined position X. On the other hand, let's assume that the right position of the slave device 30 is the initial position B of the workpiece W, and the robot arm 32 of the slave device 30 attempts to perform the task according to the learned model 12 of the first master device 10. In this case, the robot arm 22 may not be able to perform the task well. In such a case, the operator M can operate the robot arm 22 so that the robot arm 32 moves the workpiece W from the initial position B to a predetermined position X. In other words, by the intervention of the second master device 20 in a situation where the robot arm 32 of the slave device 30 is operating according to the learned model 12 of the first master device 10, the slave device 30 may be able to perform the task well.

[0018] Here, it should be noted that in the learning system 1 according to Embodiment 1, the first master device 10 is operating even when the robot arm 22 is being operated by the operator M. In other words, when operator M is operating the robot arm 22, the state of the robot arm 32 of the slave device 30 follows the state of the virtual first observed object of the first master device 10 and the state of the robot arm 22 of the second master device 20 through multilateral control. Then, the learning system 1 according to Embodiment 1 uses the observed values ​​when the slave device 30 is executing a task with the intervention of the second master device 20 to retrain the trained model 12 so that the slave device 30 can execute the task. More details will be described later.

[0019] It should be noted that if operator M does not operate the robot arm 22, the robot arm 22 of the second master device 20 will perform substantially the same operation as the robot arm 32 of the slave device 30. In other words, the state of the robot arm 22 of the second master device 20 and the state of the robot arm 32 of the slave device 30 will follow the state indicated by the first observed value output from the learned model 12 of the first master device 10.

[0020] Figure 2 shows the configuration of the learning system 1 according to Embodiment 1. The learning system 1 includes a first master device 10, a second master device 20, a slave device 30, an imaging device 50, a learning unit 60, and a control device 100. The control device 100 is communicated with the first master device 10, the second master device 20, and the slave device 30 via a wired or wireless network. The first master device 10 is also communicated with the imaging device 50 via a wired or wireless network. The control device 100 may be integrated with the first master device 10. The learning unit 60 may be provided in the first master device 10 or in the control device 100. Alternatively, the learning unit 60 may be provided in a separate device (learning device) from the other devices described above.

[0021] The first master device 10 has the functionality of a computer. The first master device 10 may be, for example, a PC (Personal Computer) or an information processing device such as a server. The first master device 10 includes a trained model 12, a command value acquisition unit 14, an inference unit 16, and an observation value output unit 18. The trained model 12 can be stored in a storage device such as memory. The command value acquisition unit 14, the inference unit 16, and the observation value output unit 18 can be realized by a control device such as a processor mounted on the first master device 10 executing a program stored in the storage device. The functions of each of the above-mentioned components of the first master device 10 will be described later.

[0022] The second master device 20 may have computer functions. The second master device 20 includes a robot arm 22, a command value acquisition unit 24, an arm control unit 26, and an observation value output unit 28. The functions of each of the above-mentioned components of the second master device 20 will be described later. The slave device 30 may have computer functions. The slave device 30 includes a robot arm 32, a command value acquisition unit 34, an arm control unit 36, and an observation value output unit 38. The functions of each of the above-mentioned components of the slave device 30 will be described later.

[0023] The imaging device 50 is, for example, a camera. The imaging device 50 captures images of the environment in which the slave device 30 is performing a task. Specifically, the imaging device 50 captures images of the environment in which the robot arm 32 of the slave device 30 is moving the workpiece W. The imaging device 50 transmits the captured images obtained through the capture to the first master device 10 and the learning unit 60.

[0024] The learning unit 60 is implemented by a device that has the functionality of a computer. The learning unit 60 generates a trained model 12 using an imitation learning method. The learning unit 60 also acquires training data that shows the slave device 30 operating using the trained model 12, and uses the training data to retrain the trained model 12. The learning unit 60 trains the trained model 12 so that the slave device 30 executes a task, using at least the command values ​​for the first master device 10 and the command values ​​for the slave device 30 as training data. The command values ​​for the first master device 10 are calculated from the second and third observed values. More specifically, the learning unit 60 trains the trained model 12 so that the slave device 30 executes a task, using the captured images obtained from the imaging device 50, the command values ​​for the first master device 10, and the command values ​​for the slave device 30 as training data. More details will be described later.

[0025] The control device 100 has the functionality of a computer. The control device 100 performs multilateral control on the first master device 10, the second master device 20, and the slave device 30. Through multilateral control, the control device 100 controls the devices so that the state of the first observed object of the first master device 10, the state of the robot arm 22 of the second master device 20, and the state of the robot arm 32 of the slave device 30 correspond to each other. In other words, the control device 100 is configured to perform coordinated control of the three devices described above. Therefore, the control device 100 functions as a coordinated control unit. Note that the functions of the control device 100 may be realized by other devices (for example, the first master device 10). As will be described later, the control device 100 uses the first observed value, the second observed value, and the third observed value to calculate command values ​​(reference values) for the first master device 10, the second master device 20, and the slave device 30, respectively.

[0026] The first master device 10 outputs a first observed value in accordance with the command value (first command value) to the first master device 10. Specifically, the first master device 10 outputs a first observed value based on the first command value and the captured image. Specifically, the command value acquisition unit 14 acquires the first command value from the control device 100. The inference unit 16 uses the trained model 12 to infer the state of the first observed object at a timing (second timing) that is a predetermined time later than the timing (first timing) corresponding to the first command value. Specifically, the inference unit 16 inputs the first command value and the captured image at the first timing to the trained model 12. Then, the inference unit 16 acquires the first observed value at the second timing output from the trained model 12. The observed value output unit 18 outputs the first observed value to the control device 100. The predetermined time corresponds to the time from when the robot (second master device 20 and slave device 30) receives a command value until the robot arms (robot arms 22, 32) enter a state corresponding to the command value.

[0027] Furthermore, the second master device 20 operates the robot arm 22 according to a command value (second command value) sent to the second master device 20. Specifically, the command value acquisition unit 24 acquires the second command value from the control device 100. The arm control unit 26 controls the operation of the robot arm 22 so that the robot arm 22 is in a state corresponding to the second command value. The arm control unit 26 may also control the joints of the robot arm 22. The observation value output unit 28 detects the state of the robot arm 22 and acquires the second observation value. The observation value output unit 28 may be implemented by, for example, a sensor. The observation value output unit 28 transmits the second observation value to the control device 100. The observation value output unit 28 may also detect the state of the robot arm 22 at predetermined intervals as described above and transmit the second observation value to the control device 100. The same applies to the slave device 30.

[0028] Furthermore, the slave device 30 operates the robot arm 32 according to a command value (third command value) directed to the slave device 30. Specifically, the command value acquisition unit 34 acquires the third command value from the control device 100. The arm control unit 36 ​​controls the operation of the robot arm 32 so that it is in a state corresponding to the third command value. The arm control unit 36 ​​may also control the joints of the robot arm 32. The observation value output unit 38 detects the state of the robot arm 32 and acquires the third observation value. The observation value output unit 38 may be implemented by, for example, a sensor. The observation value output unit 38 transmits the third observation value to the control device 100.

[0029] Figure 3 shows the configuration of the control device 100 according to Embodiment 1. The control device 100 has, as its main hardware components, a control unit 102, a storage unit 104, a communication unit 106, and an interface unit 108 (IF; Interface). The control unit 102, storage unit 104, communication unit 106, and interface unit 108 are interconnected via a data bus or the like. When the control device 100 is implemented using multiple computers, each of the multiple computers may have the hardware configuration shown in Figure 3. In addition, the other devices described above using Figure 2 (devices that implement the first master device 10, the second master device 20, the slave device 30, the imaging device 50, and the learning unit 60) may also have the hardware configuration shown in Figure 3.

[0030] The control unit 102 is a processor, such as a CPU (Central Processing Unit). The control unit 102 has the function of an arithmetic unit that performs control processing and arithmetic processing. The control unit 102 may have multiple processors. The storage unit 104 is a storage device, such as a memory or hard disk. The storage unit 104 is, for example, a ROM (Read Only Memory) or RAM (Random Access Memory). The storage unit 104 has the function of storing control programs and arithmetic programs executed by the control unit 102. In other words, the storage unit 104 (memory) stores one or more instructions. The storage unit 104 also has the function of temporarily storing processing data, etc. The storage unit 104 may include a database. The storage unit 104 may also have multiple memories.

[0031] The communication unit 106 performs the necessary processing for communicating with other devices via a network. The communication unit 106 may include a communication port, router, firewall, etc. The interface unit 108 is, for example, a user interface (UI). The interface unit 108 has an input device such as a keyboard, touch panel, or mouse, and an output device such as a display or speaker. The interface unit 108 may be configured such that the input device and the output device are integrated, for example, a touchscreen (touch panel). The interface unit 108 accepts data input operations from the user (operator) and outputs information to the user.

[0032] Furthermore, the control device 100 includes, as components, an observation value acquisition unit 120, a command value calculation unit 130, and a command value output unit 140. Each of the above-described components can be realized, for example, by executing a program under the control of the control unit 102. More specifically, each component can be realized by the control unit 102 executing a program (instruction) stored in the storage unit 104. Alternatively, each component can be realized by recording the necessary program on an arbitrary non-volatile recording medium and installing it as needed. Moreover, each component is not limited to being realized by software through a program, but may also be realized by any combination of hardware, firmware, and software. Furthermore, each component may be realized using a user-programmable integrated circuit, such as an FPGA (field-programmable gate array) or a microcontroller. In this case, the program composed of the above-described components may be realized using this integrated circuit.

[0033] The observation value acquisition unit 120 acquires observation values ​​from the first master device 10, the second master device 20, and the slave device 30. Specifically, the observation value acquisition unit 120 acquires a first observation value output from the trained model 12 from the first master device 10. The observation value acquisition unit 120 acquires a second observation value indicating the state of the robot arm 22 from the second master device 20. The observation value acquisition unit 120 acquires a third observation value indicating the state of the robot arm 32 from the slave device 30.

[0034] The command value calculation unit 130 uses the first observed value, the second observed value, and the third observed value to calculate command values ​​(reference values) for the first master device 10, the second master device 20, and the slave device 30, respectively. Specifically, the command value calculation unit 130 calculates the first command value for the first observed object of the first master device 10 based on the second observed value and the third observed value. The command value calculation unit 130 also calculates the second command value for the robot arm 22 (second observed object) of the second master device 20 based on the first observed value and the third observed value. The command value calculation unit 130 also calculates the third command value for the robot arm 32 (third observed object) of the slave device 30 based on the first observed value and the second observed value.

[0035] For example, the command value calculation unit 130 may calculate the first command value by calculating the average of the second observed value and the third observed value. Similarly, the command value calculation unit 130 may calculate the second command value by calculating the average of the first observed value and the third observed value. Furthermore, the command value calculation unit 130 may calculate the third command value by calculating the average of the first observed value and the second observed value. Here, "average" may be, for example, an arithmetic mean or a weighted average.

[0036] Here, the first observed value is q l_msr And the second observed value is q cl_msr And the third observed value is q f_msr Let's assume that the first command value is q. l_ref The second command value is q cl_ref And the third command value is q f_refLet it be so. At this time, the command value calculation unit 130 may calculate the first command value q, for example, by the following formula (1). l_ref It may be calculated.

Equation

[0037] Further, the command value calculation unit 130 may calculate the second command value q, for example, by the following formula (2). cl_ref It may be calculated.

Equation

[0038] Further, the command value calculation unit 130 may calculate the third command value q, for example, by the following formula (3). f_ref It may be calculated.

Equation

[0039] The command value output unit 140 outputs the first command value to the first master device 10. Further, the command value output unit 140 outputs the second command value to the second master device 20. Further, the command value output unit 140 outputs the third command value to the slave device 30.

[0040] FIG. 4 is a diagram illustrating a control block according to Embodiment 1. First, control regarding the first master device 10 (Leader) will be described. The control device 100 (command value calculation unit 130) uses the second observation value q cl_msr and the third observation value q f_msr at the first timing to calculate the first command value q l_ref . When the first command value q l_ref at the first timing and the captured image obtained from the imaging device 50 are input to the learned model 12 (Leaned Policy), the learned model 12 outputs the first observation value q l_msr at the second timing.

[0041] Next, the control of the second master device 20 (Co-Leader Robot) will be explained. The control device 100 (command value calculation unit 130) calculates the first observed value q at the first timing. l_msr and the third observed value q f_msr Using this, the second command value q cl_ref The arm control unit 26 (PD Controller) controls the robot arm 22 by PD (Proportional Derivative) control. Specifically, the arm control unit 26 calculates the second command value q at the first timing. cl_ref and the second observed value q cl_msr The torque command value τ used to operate the robot arm 22 is used. cl_ref The torque command value τ is calculated. cl_ref When the robot arm 22 moves, the observation value output unit 28 outputs the second observation value q at the second timing. cl_msr The following is obtained. Note that the control device 100 is set to the torque command value τ cl_ref The following calculation may also be performed. In this case, the function of the arm control unit 26 shown in Figure 4 may be implemented by the control device 100. The same applies to the slave device 30.

[0042] Next, the control of the slave device 30 (Follower Robot) will be explained. The control device 100 (command value calculation unit 130) calculates the first observed value q at the first timing. l_msr and the second observed value q cl_msr Using this, the third command value q f_ref The arm control unit 36 ​​(PD Controller) controls the robot arm 32 by PD control. Specifically, the arm control unit 36 ​​calculates the third command value q at the first timing. f_ref and the third observed value q f_msr The torque command value τ used to operate the robot arm 32 is used. f_ref The torque command value τ is calculated. f_ref When the robot arm 32 moves, the observation value output unit 38 outputs the third observation value q at the second timing. f_msrYou can obtain this.

[0043] Furthermore, as described above, the learning unit 60 receives the first command value q l_ref and the captured image and the third command value q f_ref Based on this, the trained model 12 is retrained. Specifically, the learning unit 60 acquires training data 70 which includes input data 72 (Observation) and output data 74 (Action). Here, the input data 72 is the first command value q at the first timing. l_ref This includes the captured image. Furthermore, the output data 74 includes the third command value q at the second timing. f_ref This includes the following. The learning unit 60 then retrains the trained model 12 so that it takes the input data 72 as input and outputs output data 74. As a method of retraining by imitation learning, for example, a diffusion model called Diffusion Policy (https: / / diffusion-policy.cs.columbia.edu / ) can be used. In this case, the input data 72 is the first command value q over a predetermined period. l_ref The output data 74 may also be time-series data of the captured images. Furthermore, the output data 74 is the third command value q for a predetermined period after a predetermined period of the input data 72. f_ref Time-series data is also acceptable.

[0044] Here, let's assume that in the retraining of the trained model 12, the output value of the first master device 10, that is, the first observed value q l_msr Let's consider using this as the output data. In this case, there is a risk that the operator M's actions on the second master device 20 (operator M's intervention) will not be reflected in the output data. Therefore, when retraining the trained model 12, there is a risk that the operator M's actions on the second master device 20 will not be properly reflected in the trained model 12.

[0045] In contrast, the third command value q f_ref This is the first observed value q l_msr and the second observed value q cl_msrIt is calculated by the above. Therefore, the output data 74 described above reflects the operation of worker M on the second master device 20 (worker M's intervention). That is, the first observed value q at the second timing output by the retrained model 12 l_msr This reflects the state corrected by operator intervention before relearning. Therefore, in Embodiment 1, operator M's operations on the second master device 20 can be appropriately reflected in the relearned model 12. As a result, the learning system 1 according to Embodiment 1 can appropriately realize interactive imitation learning even when an operator intervenes. Furthermore, by having an operator intervene during the operation of the system using the relearned model 12, the relearned model 12 can be efficiently relearned.

[0046] Figure 5 is a flowchart showing the learning method performed by the learning system 1 according to Embodiment 1. The control device 100 (cooperative control unit) performs cooperative control of the operation of the first master device 10, the operation of the second master device 20, and the operation of the slave device 30, as described above (step S102). The learning unit 60 performs retraining processing of the learned model 12, as described above (step S104).

[0047] (modified version) It should be noted that the present invention is not limited to the embodiments described above, and can be modified as appropriate without departing from the spirit of the invention. For example, the order of the flowchart described above can be changed as appropriate. Also, one or more of the processes in the flowchart described above can be omitted as appropriate.

[0048] The program described above, when loaded into a computer, includes a set of instructions (or software code) for causing the computer to perform one or more of the functions described in the embodiments. The program may be stored on a non-temporary computer-readable medium or a physical storage medium. Examples, but not limited to, include random-access memory (RAM), read-only memory (ROM), flash memory, solid-state drive (SSD) or other memory technologies, CD-ROM, digital versatile disk (DVD), Blu-ray® disc or other optical disc storage, magnetic cassette, magnetic tape, magnetic disk storage or other magnetic storage devices. The program may be transmitted over a temporary computer-readable medium or a communication medium. Examples, but not limited to, include electrical, optical, acoustic or other forms of propagating signals. The program may include a program product. [Explanation of symbols]

[0049] 1...Learning system, 10...First master device, 12...Trained model, 14...Command value acquisition unit, 16...Inference unit, 18...Observed value output unit, 20...Second master device, 22...Robot arm, 24...Command value acquisition unit, 26...Arm control unit, 28...Observed value output unit, 30...Slave device, 32...Robot arm, 34...Command value acquisition unit, 36...Arm control unit, 38...Observed value output unit, 50...Imaging device, 60...Learning unit, 100...Control device, 120...Observed value acquisition unit, 130...Command value calculation unit, 140...Command value output unit

Claims

1. A first master device that outputs a first observed value indicating the state of a first observed object using a pre-trained model that has been trained in advance by machine learning, A second master device having a second observation target that can be operated by an operator, A slave device having a third object to be observed, which performs a task by the operation of the third object to be observed, A coordinate control unit that controls the state of the first observed object of the first master device, the state of the second observed object of the second master device, and the state of the third observed object of the slave device to correspond to each other, A learning unit that learns the learned model so that the slave device performs the task, using a command value for the first master device calculated from a second observed value indicating the state of the second observed object of the second master device and a third observed value indicating the state of the third observed object of the slave device, and a command value for the slave device, A learning system that has the following features.

2. The aforementioned cooperative control unit, Based on the second observed value of the second master device and the third observed value of the slave device, the first command value for the first observed object of the first master device is calculated. Based on the first observed value of the first master device and the third observed value of the slave device, the second command value for the second observed object of the second master device is calculated. Based on the first observed value of the first master device and the second observed value of the second master device, a third command value is calculated for the third observed target of the slave device. The learning unit takes the first command value at least at a first timing as input and learns the trained model such that it outputs the third command value at a second timing that is a predetermined time later than the first timing. The learning system according to claim 1.

3. The aforementioned cooperative control unit, The first command value is calculated by calculating the average of the second observed value and the third observed value. The second command value is calculated by calculating the average of the first observed value and the third observed value. The third command value is calculated by calculating the average of the first observed value and the second observed value. The learning system according to claim 2.

4. The learning unit takes the first command value at the first timing and the captured image obtained by capturing the environment in which the slave device is performing the task at the first timing as inputs, and learns the trained model so as to output the third command value at the second timing. The first master device outputs the first observed value based on the first command value and the captured image obtained by capturing the environment in which the slave device operates. The learning system according to claim 2 or 3.

5. Control is performed so that the state of the first observed object in a first master device, which outputs a first observed value indicating the state of the first observed object using a pre-trained model learned by machine learning, the state of the second observed object in a second master device, which has a second observed object that can be operated by an operator, and the state of the third observed object in a slave device, which has a third observed object and performs a task by the operation of the third observed object, correspond to each other. The trained model is trained so that the slave device performs the task, using a command value for the first master device calculated from a second observed value indicating the state of the second observed object of the second master device and a third observed value indicating the state of the third observed object of the slave device, and a command value for the slave device. Learning methods.

Citation Information

Patent Citations

  • Position-force reproducing method and position-force reproducing device

    JP2009279699A