Operation learning device, operation learning system, and operation learning method for robot
The system addresses the challenge of adapting robot control to environmental changes by using shared learning models to transfer high-quality motion information across robots with different structures, reducing the learning load and enhancing adaptability.
Patent Information
- Application Number
- JP2024056408
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2024-03-29
- Publication Date
- 2025-10-10
AI Technical Summary
Existing robot control methods struggle to adapt to environmental changes and require significant learning loads, including acquiring training data and tuning parameters, especially for robots with different structures and mechanisms, limiting their applicability in dynamic environments.
A system comprising first and second learning models for individual robots, a shared learning model, and a management unit to facilitate the transfer of learning information and models across robots with varying structures and mechanisms, reducing the learning load by leveraging teacher data to train these models.
Enables the sharing of high-quality motion learning across diverse robots, reducing the need for extensive training data acquisition and parameter tuning, thereby enhancing adaptability to environmental changes.
Smart Images

Figure 2025153776000001_ABST
Abstract
Description
[Technical Field]
[0001] The present invention relates to a robot motion learning device, a motion learning system, and a motion learning method that enable the sharing or transfer of learning information and learning models between multiple types of robots, particularly between robots with different mechanisms, structures, characteristics, etc., in a robot system that autonomously generates robot motion control sequences by learning robot motion data. [Background technology]
[0002] Work at manufacturing and construction sites, and maintenance work on infrastructure facilities such as railways, plants, power plants, and buildings require skilled personnel and is dangerous and heavy work, making it difficult to secure workers for these tasks, and there are high hopes for automation using robots. However, conventional control methods, in which all of a robot's movements are written as programs, cannot respond to situations that are not described. As a result, the scope of robot use is limited to applications where the environment must be maintained to remain constant and the same tasks must be repeated, making it difficult to apply robots to tasks that require adaptation to environmental changes such as those described above.
[0003] Therefore, artificial intelligence (AI) that uses learning techniques such as neural computing is attracting attention. For example, by utilizing deep learning, it is possible to deal with a certain degree of environmental change using the generalization capabilities of deep learning without writing a program. Furthermore, even in the case of major environmental changes, by learning training data for operating a robot in that environment, it will be possible to adapt to new situations.
[0004] However, to fully utilize the advantages of these learning methods, it is necessary to properly prepare the training data used for learning and to properly assign parameters to the motion learning model that acquires motion information. Insufficient training data and parameters will prevent the acquisition of robust, generalizable motions. Training data requires motion data from multiple situations to adapt to environmental changes. Training data is prepared by combining data from simulations and data from actual robot operations. Reinforcement learning, a deep learning method, requires tens of thousands to hundreds of millions of training data sets. Thus, acquiring training data is a significant challenge for these learning methods. Therefore, if a motion learning model that has acquired high-quality, robust, and generalizable motions could be utilized for other robots, the burden of acquiring training data, such as the overhead of acquiring training data, could be reduced.
[0005] Patent Document 1 discloses a method for obtaining a general-purpose trained model by integrating multiple individually trained models obtained by training based on individual motion data acquired by a group of motion devices having the same configuration. [Prior art documents] [Patent documents]
[0006] [Patent Document 1] Japanese Patent Publication No. 2023-89023 Summary of the Invention [Problem to be solved by the invention]
[0007] The learning method described in Patent Document 1 can autonomously generate robust behaviors even in the face of environmental changes, but in order to acquire high-quality behaviors, the learning load, such as obtaining training data of appropriate quality and quantity, parameter tuning, and computational costs, is an issue. When using multiple robots with the same structure, the learning load can be reduced by using multiple robots to collect and integrate training data of different movements, as described in Patent Document 1. Robots with the same structure refer to a range of structures and characteristics that can be considered identical.
[0008] Furthermore, if multiple robots have the same structure but different characteristics, such as the correction amount for the target stopping position, and the correction amount is clear and numerically correlated to the differences in characteristics between the robots, the learning load can be reduced by adding the correction amount to the method described in Patent Document 1 and transferring the learning results of the trained robot to an untrained robot. However, if the differences in characteristics are unknown even for robots with the same structure, or if the robots have different mechanisms and structures, the trained learning information and learning model cannot be shared or transferred. Therefore, appropriate training data must be acquired for each robot and parameters must be tuned. In other words, the learning load required to achieve high-quality motion is an issue. As a result, the invention described in Patent Document 1 has the problem of being unable to achieve high-quality motion. This issue is particularly pronounced in the case of robots used in manufacturing and construction sites, or for maintenance of infrastructure such as railways, plants, power plants, and buildings, where various types of robots exist depending on the location.
[0009] Therefore, an object of the present invention is to reduce the learning load on a learning model that controls multiple types of robots. [Means for solving the problem]
[0010] In order to solve the above-mentioned problems, the robot behavior learning device of the present invention is characterized by comprising: a plurality of first learning models that, for each of a plurality of types of robots, accept behavior information at a certain time and convert it into behavior features, and accept external world information at that time and convert it into external world features; a shared learning model that converts the behavior features and external world features output by each of the first learning models into predicted behavior features for the next time that are common to a plurality of types of robots; a plurality of second learning models that, for each of the plurality of robots, convert the predicted behavior features for the next time into predicted behavior information; and a management unit that uses teacher data regarding the behavior of each of the robots to train any of the first learning model, the second learning model, and the shared learning model for that robot.
[0011] The robot movement learning system of the present invention is characterized by including the robot movement learning device described above and a plurality of types of robots.
[0012] The robot movement learning method of the present invention is a robot movement learning method for learning the movements of multiple types of robots, and is characterized by including the steps of: a first learning model corresponding to a certain robot using training data regarding the movement of the robot to learn a process for converting movement information and external world information of the robot at a certain time into common movement features; a shared learning model using the training data regarding the movements of the multiple types of robots to learn the time series relationship of common movement features regarding movements common to the multiple types of robots; and a second learning model corresponding to the robot using the training data regarding the movement of the robot to learn a process for converting a predicted value at the next time of the common movement features output by the shared learning model into predicted movement information of the robot at the next time. Other means will be described in the detailed description of the invention. [Effects of the Invention]
[0013] According to the present invention, it is possible to reduce the learning load on a learning model that controls multiple types of robots. [Brief explanation of the drawings]
[0014] [Figure 1] 1 is a diagram illustrating an example of the configuration of a robot action learning device according to a first embodiment. [Figure 2] 1 is an example of the configuration of a motion learning system. [Figure 3] 1 is an example of a hardware configuration of an action learning device. [Figure 4] 1 is an example of a hardware configuration of an action learning device. [Figure 5] FIG. 2 is a diagram illustrating an example of a detailed configuration of an action learning device. [Figure 6] FIG. 2 is a diagram illustrating an example of a detailed configuration of an action learning device. [Figure 7] FIG. 2 is a diagram illustrating an example of a detailed configuration of an action learning device. [Figure 8] 1 shows the overall sequence of learning for the action learning device. [Figure 9A] This is an initial learning process of the action learning device. [Figure 9B] This is an initial learning process of the action learning device. [Figure 9C] This is an initial learning process of the action learning device. [Figure 9D] This is an initial learning process of the action learning device. [Figure 10A] 10 shows an unlearned action learning process of the action learning device. [Figure 10B] 10 shows an unlearned action learning process of the action learning device. [Figure 10C] 10 shows an unlearned action learning process of the action learning device. [Figure 11A] This is a process for adding a new robot to the motion learning device. [Figure 11B] This is a process for adding a new robot to the motion learning device. [Figure 11C] This is a process for adding a new robot to the motion learning device. [Figure 12A]1 shows an individual robot learning process of the motion learning device. [Figure 12B] 1 shows an individual robot learning process of the motion learning device. [Figure 12C] 1 shows an individual robot learning process of the motion learning device. [Figure 12D] 1 shows an individual robot learning process of the motion learning device. DETAILED DESCRIPTION OF THE INVENTION
[0015] Hereinafter, embodiments for carrying out the present invention will be described in detail with reference to the drawings. Note that the embodiments described below are examples of means for realizing the present disclosure, and should be appropriately modified or changed depending on the configuration of the device to which the present disclosure is applied and various conditions, and the present disclosure is not limited to the following embodiments.
[0016] System Configuration FIG. 1 is a diagram showing an example of the configuration of a robot action learning device 1 according to this embodiment. In this embodiment, the system includes a plurality of types of robots 2a to 2c, a plurality of first learning models 3a to 3c, a plurality of second learning models 4a to 4c, a shared learning model 5, a management unit 51, and an action designation unit 52.
[0017] The first learning models 3a to 3c receive motion information at a certain time for each of the multiple types of robots 2a to 2c and convert it into motion feature quantities, and also receive external world information at this time and convert it into external world feature quantities.
[0018] The shared learning model 5 converts the motion features output by the first learning models 3a to 3c into predicted motion features for the next time that are common to multiple types of robots 2a to 2c, and also converts the external environment features output by the first learning models 3a to 3c into predicted external environment features for the next time that are common to multiple types of robots 2a to 2c.
[0019] The plurality of second learning models 4a to 4c convert the predicted motion feature amount for the next time output by the shared learning model 5 into predicted motion information for the next time for each of the plurality of types of robots 2a to 2c. The management unit 51 uses teacher data relating to the movements of each of the robots 2a to 2c to train one of the first learning model, second learning model, and shared learning model 5 relating to one of the robots. When a plurality of actions have been learned, the action designation unit 52 selects a desired action from the action learning device 1 and causes each of the robots 2a to 2c to execute the action.
[0020] The robots 2a to 2c are robots that wish to share the desired motion acquired by the shared learning model 5, and there may be any number of them. The desired motion is, for example, a task or a series of tasks, such as grasping an object within the field of view or opening and closing a door. This task may be, but is not limited to, a task related to manufacturing, maintenance, or housework, such as attaching parts, welding, painting, or drilling.
[0021] The first learning models 3a to 3c correspond to the robots 2a to 2c, respectively, and receive input of information on the sensors and robot status installed in the robots 2a to 2c. The sensors include image and distance sensors using imaging elements and lasers, sensors for measuring forces acting on various parts of the robots 2a to 2c, and tactile sensors for measuring contact with objects. Information on the status of the robots 2a to 2c includes the joint angles of the robots 2a to 2c and motor current values. The first learning models 3a to 3c receive input of this information for each robot 2a to 2c, learn and extract features related to the robot's operation from information on the external world (sensors) and the internal world (robot status), and output the resulting features to the shared learning model 5.
[0022] The shared learning model 5 is located between the multiple first learning models 3a-3c and the multiple second learning models 4a-4c. Regardless of the number of robots 2a-2c, one shared learning model 5 is required for each desired motion or task to be shared among the robots 2a-2c. The shared learning model 5 learns a sequence of shared motions or tasks. The shared learning model 5 outputs future motion feature values to be transitioned from the current motion feature values input from the first learning models 3a-3c, and inputs these to the multiple second learning models 4a-4c. Here, the future motion feature values to be transitioned are basically the motion feature values for the next time in the robot's control cycle. If the control cycles of the robots 2a-2c differ, the control cycles of the robots 2a-2c are adjusted by interpolation, synchronization, or the like. To calculate the future motion feature values to be transitioned, the shortest control cycle among the robots 2a-2c may be used.
[0023] The second learning models 4a to 4c learn the relationship between the future motion feature values to be changed that are input from the shared learning model 5 and the motions and control outputs of the corresponding robots 2a to 2c at that time, and output the resulting control outputs to the robots 2a to 2c. As a result, each robot 2a to 2c executes a series of sequences of desired motions or tasks.
[0024] <Hardware Configuration> FIG. 2 is a diagram showing an example of the configuration of an action learning system 60 in this embodiment. The motion learning system 60 of this embodiment shown in Fig. 2 includes a motion learning device 1, multiple types of robots 2a to 2d, a network 61, and a robot motion teaching device 64. Examples of the multiple types of robots 2a to 2d are assumed here to include robots that work on construction sites, robots that work on manufacturing sites, and robots that perform household chores at home, but the present invention is not limited to these. The motion learning device 1 generates and saves a motion model that shares a common motion among these robots 2a to 2d, such as the motion of grasping an object within the camera's field of view, and is capable of transferring a new shared motion learned by one robot to the motion of another robot.
[0025] The network 61 may be the Internet, a telephone network, or the like. The action learning device 1 is an information processing device that stores, for example, parameter and weight information, action data of each robot, and teacher data. Here, weight information refers to, for example, the weight between network elements of a learning model. The action learning device 1 operates in cooperation with a cloud server, a hard disk connected to a LAN (Local Area Network), and the like. The multiple types of robots 2a to 2d, the robot action teaching device 64, and the action learning device 1 are configured to be able to access each other as appropriate.
[0026] The action learning device 1 interfaces with an action learning administrator, accesses necessary information through communication with the robots 2a to 2d and a server via a network 61, and executes learning of a learning model. Note that the calculation itself may be performed using a server connected to the network 61 (not shown), and is not limited thereto.
[0027] The robot motion teaching device 64 is one form of means for acquiring motion data of each of the robots 2a to 2d. The motion data of each robot includes external world information detected by sensors in each of the robots 2a to 2d and internal world information indicating the state of each of the robots 2a to 2d. The robot motion teaching device 64 is, for example, a remote control device in which a person remotely controls the robots 2a to 2d using an AR (Augmented Reality) system that uses camera images mounted on the robots 2a to 2d, or a haptics system that presents reaction forces and tactile sensations acting on the robots 2a to 2d.
[0028] 3 is a diagram showing an example of the hardware configuration of the robots 2a to 2d in this embodiment. Hereinafter, when there is no need to particularly distinguish between the robots 2a to 2d, they will be simply referred to as robot 2.
[0029] The robot 2 includes a processing unit 70, a communication interface 77, a display unit 75, and an input unit 76. The calculation processing unit 70 includes a CPU 71, a ROM 72, a RAM 73, an external memory 74, and a system bus 78. The communication interface 77 is an interface with the network 61. The display unit 75 and the input unit 76 are interfaces with the administrator. The calculation processing unit 70 executes a predetermined machine learning program and sets the configuration and parameters of the behavior model downloaded from the behavior learning device 1, thereby realizing the first learning model, the shared learning model, and the second learning model.
[0030] The CPU 71 comprehensively executes information processing in the arithmetic processing unit 70 and controls the other components via a system bus 78. The ROM 72 is a non-volatile memory that stores control programs and the like required for the CPU 71 to execute processing. The programs may be stored in the external memory 74 or a removable storage medium. The RAM 73 is a volatile memory that operates as the main memory of the CPU 71 and functions as a work area, etc. In other words, when executing processing, the CPU 71 loads necessary programs and data from the ROM 72 or the external memory 74 into the RAM 73 and executes the programs to perform various functional operations.
[0031] The external memory 74 can store various data and information required when the CPU 71 performs processing using a program, as well as intermediate and final processing results. The external memory 74 stores parameter and weight information, the robot's own operation data and training data, the program that realizes the processing, the robot's own status, and the like. The weight information is, for example, the weight between network elements of a learning model.
[0032] The display unit 75 is configured by a monitor such as a liquid crystal display, etc. The input unit 76 is configured so that the manager of the robot 2 can give instructions to the robot 2.
[0033] The communication interface 77 is an interface for communicating with an external device. In this embodiment, the communication interface 77 communicates with the motion learning device 1, the robot motion teaching device 64, etc. The communication interface 77 can be, for example, a wireless communication LAN (Local Area Network) interface or a wired communication LAN interface. The system bus 78 connects the CPU 71, ROM 72, RAM 73, external memory 74, display unit 75, input unit 76, communication interface 77, external world / internal world measuring unit 80, and actuator 81 so as to enable communication between them.
[0034] The external and internal world measurement unit 80 is composed of various sensors. The robot 2's external sensors include, for example, image sensors and distance image sensors using image sensors or lasers, sensors measuring forces and torques acting on various parts of the robot 2, and tactile sensors measuring the proximity and contact states between the robot 2 and objects. The robot 2's internal sensors include angle sensors that measure the robot 2's joint angles and motor voltage and current sensors. Each robot 2 has different configurations, performance, and output formats. Therefore, the present invention is configured to enable the sharing and transfer of learning information and learning models between such different robots. High-quality movements acquired by one robot can be shared and transferred to other robots that have not yet learned them. This reduces the learning load, such as acquiring training data and parameter tuning, and facilitates the acquisition of high-quality movements. The external and internal world measurement unit 80 is also an essential component of the robot motion teaching device 64. The robot motion teaching device 64 measures the angles and pressures of the robot's onboard switches and movable mechanisms, and controls the robot 2 based on this information.
[0035] The actuator 81 consists of actuators that move hardware mechanisms and electronic components that control the actuator output. Actuators 81 include motors, which are rotating elements that use electromagnetic force; solenoids, which are linear-motion elements; and vibration elements such as piezos. In the robot 2, actuators are used for wheels, arm joints, hand opening and closing, and camera pan-tilt. Robots 2 have various mechanisms, such as the number of joints, arm length, and number of fingers. The present invention enables the sharing and transfer of learning information and learning models between different robots, allowing high-quality motions acquired by one robot to be shared and transferred to other unlearned robots. This reduces the learning load, such as acquiring training data and parameter tuning, and facilitates the acquisition of high-quality motions. The robot motion teaching device 64 also includes actuators when presenting reaction forces and tactile sensations to the robot 2.
[0036] FIG. 4 is a diagram showing an example of the hardware configuration of the action learning device 1 in this embodiment. The action learning device 1 includes a calculation processing unit 70, a communication interface 77, a display unit 75, and an input unit 76. The arithmetic processing unit 70 includes a CPU 71, a ROM 72, a RAM 73, an external memory 74, and a system bus 78. A communication interface 77 is an interface with the network 61. A display unit 75 and an input unit 76 are interfaces with the administrator.
[0037] The CPU 71 comprehensively executes information processing in the arithmetic processing unit 70 and controls the other components via a system bus 78. The ROM 72 is a non-volatile memory that stores control programs and the like required for the CPU 71 to execute processing. The programs may be stored in the external memory 74 or a removable storage medium. The RAM 73 is a volatile memory that operates as the main memory of the CPU 71 and functions as a work area, etc. In other words, when executing processing, the CPU 71 loads necessary programs and data from the ROM 72 or the external memory 74 into the RAM 73 and executes the programs to perform various functional operations.
[0038] The external memory 74 can store various data and information required when the CPU 71 performs processing using a program, as well as intermediate and final processing results. The external memory 74 stores configuration information for the action learning device 1, parameter and weight information, action data and teacher data for each of the robots 2a to 2d, a program for implementing the processing, and the status of each of the robots 2a to 2d. The weight information is, for example, the weight between network elements of a learning model.
[0039] The display unit 75 is configured with a monitor such as a liquid crystal display. The input unit 76 is configured with a pointing device such as a keyboard and a mouse. The input unit 76 is configured so that the administrator can check information from each device and give instructions.
[0040] The communication interface 77 is an interface for communicating with external devices. In this embodiment, the communication interface 77 communicates with multiple types of robots 2a to 2c, a robot operation teaching device 64, etc. The communication interface 77 can be, for example, a wireless communication LAN (Local Area Network) interface or a wired communication LAN interface. The system bus 78 connects the CPU 71, ROM 72, RAM 73, external memory 74, display unit 75, input unit 76, and communication interface 77 so as to enable communication between them.
[0041] <<An Example of Detailed Configuration of Action Learning Device 1>> An embodiment of the detailed configuration of the action learning device 1 will be described below with reference to Fig. 5 to Fig. 7. Fig. 5 to Fig. 7 are diagrams for explaining the inside of the system configuration example of Fig. 1 in more detail.
[0042] FIG. 5 shows an example of the detailed configuration and operation of the first learning model 3a and second learning model 4a for the robot 2a, and the shared learning models 5a and 5b that share the learning results of the robots 2a to 2d.
[0043] The first learning model 3a includes machine learning models 31 and 32. The machine learning models 31 and 32 include, for example, a convolutional neural network (CNN) or an autoencoder (AE). The autoencoder (AE) included in the machine learning models 31 and 32 is a configuration that reduces information in multiple fully connected layers while reducing the number of elements in a neural network. The first learning model 3a of the robot 2a receives image information i, which is external world information, from a camera. t and the robot motion information a, which is internal information, is obtained from the joint angle sensors. t Accepts input.
[0044] The machine learning model 31 uses external information, i.e., image information t The machine learning model 32 extracts the external feature 91a from the internal information, i.e., the robot motion information a t The internal boundary feature 92a is extracted from the
[0045] The shared learning models 5a and 5b include a learning model 53 that learns time-series information in which a recursive loop 99 exists, such as an RNN (Recurrent Neural Network) or an LSTM (Long Short Term Memory). The learning model 53 acquires an action sequence and an action model. Here, an example of an RNN is shown.
[0046] The learning model 53, which is an RNN, outputs features for a future time (t+1) to be transitioned from features for the current time t. The external world features 91a output from the first learning model 3 are connected one-to-one to the external world features 93 on the input side of the shared learning models 5a and 5b. The internal world features 92a output from the first learning model 3 are connected one-to-one to the internal world features 94 on the input side of the shared learning models 5a and 5b. Since multiple types of robots 2a do not use the same action learning device 1 simultaneously in either the learning process or the execution process, this connection is switched on a robot-by-robot basis. Note that the first learning model 3a and the second learning model 4a are learning models specific to this robot 2a.
[0047] The predicted external environment feature 95, which is the output side of the shared learning models 5a and 5b, is connected one-to-one to the predicted external environment feature 97a of the robot 2a. The predicted internal environment feature 96, which is the output side of the shared learning models 5a and 5b, is connected one-to-one to the predicted internal environment feature 98a of the robot 2a.
[0048] The second learning model 4 includes machine learning models 41 and 42. The machine learning model 41 is configured with multi-layer connections while increasing the number of elements of the neural network, which is the opposite of the machine learning model 31. The machine learning model 41 calculates external information (i t+1 The machine learning model 42 generates a predicted value of the robot operation information (a t+1 ) is generated. Among the outputs of the second learning model 4, the information related to the robot's motion is the robot motion information (a t+1 ), but from the viewpoint of learning, external information (i t+1 ) predicted values are also output.
[0049] The motion designation unit 52 selects one of the shared learning models 5a, 5b by assigning a value corresponding to one of the motions to a parametric bias node that has been trained to have one value for one motion. When the motion designation unit 52 assigns a value corresponding to the desired motion to the parametric bias node, it becomes possible to select the desired motion from the motion learning device 1 that has learned multiple motions and execute the motion of the robot 2.
[0050] FIG. 6 shows an example of the detailed configuration and operation of the first learning model 3b and second learning model 4b for the robot 2b, and the shared learning models 5a and 5b that share the learning results of the robots 2a to 2d.
[0051] The first learning model 3b includes machine learning models 33 and 32. The machine learning model 33 of the first learning model 3b of the robot 2b receives image information i from the camera. t In addition to this, external information such as haptic information f t The machine learning model 33 receives input of image information i t and haptic information f t When input, the external feature 91b is output via the fully connected layer.
[0052] Robot motion information a input to machine learning model 32 t Although not shown, the robot motion information a includes information from a motor current sensor and the like in addition to the joint angles of the robot 2b. t When input, it outputs the internal boundary feature 92b.
[0053] The machine learning model 33 uses external information, i.e., image information t and haptic information f t The machine learning model 32 extracts the external feature 91b from the internal information, i.e., the robot motion information a t The internal boundary feature 92b is extracted from the
[0054] The shared learning models 5a and 5b include a learning model 53 that learns time-series information in which a recursive loop 99 exists, such as an RNN or LSTM. The learning model 53 acquires an action sequence and an action model. Here, an example of an RNN is shown.
[0055] The learning model 53, which is an RNN, outputs features for a future time (t+1) to be transitioned from features for the current time t. The external world feature 91b output from the first learning model 3 is connected one-to-one to the external world feature 93 on the input side of the shared learning models 5a and 5b. The internal world feature 92b output from the first learning model 3 is connected one-to-one to the internal world feature 94 on the input side of the shared learning models 5a and 5b. Since multiple types of robots 2b do not use the same action learning device 1 simultaneously in either the learning process or the execution process, this connection is switched on a robot-by-robot basis. Note that the first learning model 3b and the second learning model 4b are learning models specific to this robot 2b.
[0056] The predicted external environment feature 95, which is the output side of the shared learning models 5a and 5b, is connected one-to-one to the predicted external environment feature 97b of the robot 2b. The predicted internal environment feature 96, which is the output side of the shared learning models 5a and 5b, is connected one-to-one to the predicted internal environment feature 98b of the robot 2b.
[0057] The second learning model 4b includes machine learning models 43 and 42. The machine learning model 43 is configured with a multi-layer connection while increasing the number of elements of the neural network, which is the opposite of the machine learning model 33. The machine learning model 43 derives image information (i t+1 ) and haptic information (f t+1 The machine learning model 42 generates a predicted value of the robot operation information (a) at the next time from the predicted inner bound feature 98c. t+1 ) is generated. Among the outputs of the second learning model 4b, information related to the behavior of the robot 2b is the robot behavior information (a t+1 ), but from the viewpoint of learning, external information, such as image information (i t+1 ) and haptic information (f t+1) predicted values are also output.
[0058] FIG. 7 shows an example of the detailed configuration and operation of the first learning model 3c and second learning model 4c for the robot 2c, and the shared learning models 5a and 5b that share the learning results of the robots 2a to 2d.
[0059] The first learning model 3c includes machine learning models 33 and 32. The machine learning model 33 of the first learning model 3c of the robot 2c receives image information i from the camera. t In addition to this, external information such as haptic information f t The machine learning model 33 receives input of image information i t and haptic information f t When input, the external feature 91c is output via the fully connected layer.
[0060] Robot motion information a input to machine learning model 32 t Although not shown, the robot motion information a includes information from a motor current sensor and the like in addition to the joint angles of the robot 2c. t When input, it outputs the internal boundary feature 92c.
[0061] The machine learning model 33 uses external information, i.e., image information t and haptic information f t The machine learning model 32 extracts the external feature 91c from the internal information, i.e., the robot motion information a t The internal boundary feature 92c is extracted from the
[0062] The shared learning models 5a and 5b include a learning model 53 that learns time-series information in which a recursive loop 99 exists, such as an RNN or an LSTM. The learning model 53 acquires an action sequence and an action model. Here, an example of an RNN is shown.
[0063] The learning model 53, which is an RNN, outputs features for a future time (t+1) to be transitioned from features for the current time t. The external world feature 91c output from the first learning model 3 is connected one-to-one to the external world feature 93 on the input side of the shared learning models 5a and 5b. The internal world feature 92c output from the first learning model 3 is connected one-to-one to the internal world feature 94 on the input side of the shared learning models 5a and 5b. Since multiple types of robots 2c do not use the same action learning device 1 simultaneously in either the learning process or the execution process, this connection is switched on a robot-by-robot basis. Note that the first learning model 3c and the second learning model 4c are learning models specific to this robot 2c.
[0064] The predicted external environment feature 95, which is the output side of the shared learning models 5a and 5b, is connected one-to-one to the predicted external environment feature 97c of the robot 2c. The predicted internal environment feature 96, which is the output side of the shared learning models 5a and 5b, is connected one-to-one to the predicted internal environment feature 98c of the robot 2c.
[0065] The second learning model 4c includes machine learning models 43 and 42. The machine learning model 43 is configured with a multi-layer connection while increasing the number of elements of the neural network, which is the opposite of the machine learning model 33. The machine learning model 43 calculates image information (i t+1 ) and haptic information (f t+1 The machine learning model 42 generates a predicted value of the robot operation information (a) at the next time from the predicted inner bound feature 98c. t+1 ) is generated. Among the outputs of the second learning model 4c, information related to the behavior of the robot 2c is the robot behavior information (a t+1 ), but from the viewpoint of learning, external information, such as image information (i t+1 ) and haptic information (f t+1 ) predicted values are also output.
[0066] Here, the external boundary features 91a to 91c shown in Figures 5 to 7 have the same number of neuron elements so that they have the same amount of information. The internal boundary features 92a to 92c also have the same number of neuron elements so that they have the same amount of information. With this configuration, even if the configuration of the robot 2 is different, the same information can be obtained at the same scene in the work sequence.
[0067] Overall learning process FIG. 8 is a flowchart of the learning process of the action learning device 1. First, the management unit 51 executes an initial learning process for the robots 2a to 2c to obtain an initial operation model (step S11).
[0068] Next, in step S12, the management unit 51 branches and moves to one of four steps depending on the selected update condition. If an action addition is requested in step S12, the management unit 51 executes the learning process for the unlearned action (step S13), and the process returns to step S12.
[0069] If a request to add a robot is made in step S12, the management unit 51 executes a process to add a new robot (step S14), and then the process returns to step S12. If individual tuning is requested for each robot in step S12, the management unit 51 executes learning processing for each individual robot (step S15), and then the processing returns to step S12.
[0070] When the management unit 51 detects a combination of robots other than the robots 2a to 2c, or a significant change in the configuration of multiple types of robots, the process returns to step S11, and the initial learning process for creating a new model is executed (step S11).
[0071] <<Initial learning process>> The initial learning process of the action learning device 1 will be described below with reference to FIGS. 9A to 9D. The initial learning process in Figure 9A is a process of teaching the robot a desired shared behavior, such as "grabbing a part," to acquire an initial behavior model. Here, the shared behavior is a basic behavior with a low level of difficulty. The behavior model is a learning model after learning that has acquired the desired behavior, that is, a learning model that can be used when the robot 2 executes the desired behavior.
[0072] The management unit 51 determines whether or not there is an unlearned robot (step S20). If there is no unlearned robot (No), the processing of Fig. 9A ends. If there is an unlearned robot (Yes), the processing proceeds to step S21.
[0073] First, in step S21, the management unit 51 causes a target unlearned robot to perform a desired shared action, such as grasping a predetermined part, and acquires motion data related to the shared action of the robot. Here, the desired shared action is a basic action or an action with a low level of difficulty, and the same action is executed by all robots. The management unit 51 acquires external world information of the robot by detecting it with a sensor as the robot's motion data, and further acquires robot state information, which is internal world information of the robot.
[0074] The management unit 51 acquires the robot's operation data by various methods, such as remotely controlling the robot 2 using a remote control device, direct teaching, where a person holds the robot directly and performs the operation, programming the operation in advance and playing it back, or using a combination of simulations.
[0075] The management unit 51 generates teacher data for training the learning model of the action learning device 1 from the action data acquired in step S21 (step S22). The teacher data is used in the learning process of changing parameters in the learning model so as to reduce the error based on an evaluation function using the error between the teacher data and the output of the learning model during learning by the learning model of the action learning device 1. The parameters in the learning model are, for example, weights between network elements of the learning model.
[0076] Part of the training data is evaluation data used in the process of determining the convergence of learning by the learning model. Convergence of learning means that the error between the training data and the output of the learning model is below a predetermined value. For all acquired motion data, the training data is a set of input training data, which is motion data at a certain time, and output training data, which is motion data at the next time in the control cycle. If the working time or control cycle differs between the robots, the management unit 51 normalizes the working time and adjusts (interpolates, synchronizes) the cycle, and generates training data.
[0077] The management unit 51 uses the teacher data generated in step S22 to train the learning model of the action learning device 1, thereby generating an action model (step S23). The management unit 51 provides input training data to the first learning models 3a to 3c corresponding to each of the robots 2a to 2c. The management unit 51 changes the parameters of the learning models corresponding to each of the robots 2a to 2c using the error between the output of the second learning models 4a to 4c corresponding to each of the robots 2a to 2c at that time and the output training data. The parameters of the learning models are, for example, the weights between the network elements of the learning models.
[0078] In FIG. 9B, the first learning model 3a and the second learning model 4a corresponding to the robot 2a, and the shared learning model 5 that simultaneously undergoes learning, are shown by hatching. In FIG. 9C, the first learning model 3b and the second learning model 4b corresponding to the robot 2b, and the shared learning model 5 that simultaneously undergoes learning, are shown by hatching. In FIG. 9D, the first learning model 3c and the second learning model 4c corresponding to the robot 2c, and the shared learning model 5 with which learning is performed simultaneously, are shown by hatching. These processes are repeated using all the teaching data acquired by all the robots 2a to 2c until the learning converges, and as a result, a behavior model is generated.
[0079] 9A, the explanation will be continued. The management unit 51 saves the behavior model generated in step S23 (step S24). The data saved by the management unit 51 as the behavior model includes configuration information constituting the learning model of the behavior learning device 1, which will be described later, as well as parameters and weight information.
[0080] The management unit 51 stores the first learning model 3a and the second learning model 4a shown in FIG. 9B, and the shared learning model 5 that is simultaneously learned, in the robot 2a as operation models. The management unit 51 stores the first learning model 3b and the second learning model 4b shown in FIG. 9C, and the shared learning model 5 that is simultaneously learned, in the robot 2b as operation models. The management unit 51 stores the first learning model 3c and the second learning model 4c shown in FIG. 9D, and the shared learning model 5 that is simultaneously learned, in the robot 2c as operation models.
[0081] The configurations of the robots 2a to 2c may differ, for example, in the robot mechanisms, image sensors, and force-tactile sensors, or may have or not have sensors. In this way, by learning with a large number of robots / various types of robots, common information for desired operations is learned and accumulated in the shared learning model 5. Input information processing specific to each robot is automatically separated, learned, and accumulated in the first learning model 3. Output information processing specific to each robot is automatically separated, learned, and accumulated in the second learning model 4.
[0082] <Learning process for unlearned actions> The learning process of an unlearned action by the action learning device 1 will be described below with reference to Figures 10A to 10C. The management unit 51 uses training data related to actions that have not been learned by the robots 2a to 2c to train the shared learning model 5 while keeping the parameters of the first learning models 3a to 3c and the second learning models 4a to 4c fixed. This reduces the learning load on the learning models.
[0083] The basic flow of the learning process for unlearned behavior in Figure 10A is the same as the initial learning process in Figure 9A, but the part that is learned is different. For simplicity, a case will be described here in which a behavior model acquired by robot 2a is transferred to unlearned robot 2c. This combination is for convenience, and the behavior model acquired by robot 2a may also be transferred to unlearned robots 2b and 2c, or the model acquired by robots 2a and 2b may also be transferred to unlearned robot 2c.
[0084] First, the management unit 51 executes a desired unlearned shared action using, for example, the robot 2a, and acquires the action data of the robot 2a (step S31). The action data of the robot 2a is acquired by detecting external information of the robot 2a using a sensor and acquiring the robot state, which is internal information of the robot 2a.
[0085] In the initial learning process, movement data was acquired for all robots 2a to 2c to be used, but in the learning process for unlearned movements, movement data for all robots 2a to 2c is not required; it is sufficient to acquire movement data for one or more types of robots that can be executed relatively easily or that can acquire high-quality movements.
[0086] The management unit 51 generates teacher data for learning the learning model of the action learning device 1 from the action data acquired in step S31 (step S32). The process of generating this teacher data is the same as step S22 of the initial learning process.
[0087] The management unit 51 uses the training data generated in step S32 to train the action learning device 1 and generate an action model (step S33). The management unit 51 provides input training data to the first learning model 3a corresponding to the robot 2a, and changes the parameters of the shared learning model 5 shown in the hatched area in Figure 10B using the error between the output of the second learning model 4a corresponding to the robot 2a and the output training data. At the same time, the management unit 51 fixes the parameters of the first learning model 3a and the second learning model 4a.
[0088] In the initial learning process, the management unit 51 trains the learning models of the hatched areas corresponding to each of the robots 2a to 2c, as shown in Figures 9B to 9D. At the time of the learning process for an unlearned motion, the management unit 51 trains the first learning models 3a to 3c and the second learning models 4a to 4c for each of the robots 2a to 2c. Therefore, for a new motion, it is only necessary to train the robots for the series of shared motion sequences. Therefore, in step S33, the management unit 51 trains only the shared learning model 5.
[0089] The management unit 51 stores the shared learning model 5 generated in step S33 as an operation model in the robot (step S34). Specifically, the management unit 51 stores the shared learning model 5 as an operation model in the robot 2c in order to have the robot 2c perform this unlearned operation. As a result, the robot 2c can perform the unlearned operation that the robot 2c has not performed before, using the newly stored shared learning model 5, as well as the first learning model 3c and the second learning model 4c, in the configuration shown in FIG. 10C. At this time, sensing data from the external / internal world measurement unit 80 of the robot 2c is input to the first learning model 3c, and the actuator 81 of the robot 2c is driven based on the predicted operation information output via the shared learning model 5 and the second learning model 4c.
[0090] According to the learning process for unlearned movements, movements using a robot that can easily acquire a desired movement, or that can acquire a high-quality movement, or movements of a robot that has by chance acquired a high-quality movement, can be transferred to another robot, making it possible to easily acquire a high-quality movement.
[0091] Furthermore, since the management unit 51 re-learns the shared learning model 5 in this learning process, the one used in the initial learning may be used as the initial value for re-learning, or the initialized one may be used for new learning. By learning and accumulating a new shared learning model 5 for each desired action, multiple behavior models for each corresponding action can be obtained, and the desired action can be realized by selecting a behavior model at the time of execution.
[0092] <<Adding a new robot>> The new robot addition process of the action learning device 1 will be described below with reference to FIGS. 11A to 11C.
[0093] The management unit 51 newly adds a first learning model 3d and a second learning model 4d corresponding to the robot 2d to be newly added as a control target. Then, the management unit 51 uses the training data of the newly added robot 2d to train the first learning model 3d and the second learning model 4d with respect to the movements that the shared learning model 5 has already learned, while keeping the parameters of the shared learning model 5 fixed.
[0094] The basic flow of the new robot addition process in Fig. 11A is the same as the initial learning process in Fig. 9A, but the learning part is different. Here, as shown in Fig. 11B, the management unit 51 trains a first learning model 3d and a second learning model 4d related to the newly added robot 2d. Even with existing robots 2a to 2c, if the configuration changes due to sensor replacement or hand replacement, it is recommended to treat them as new robots.
[0095] First, the management unit 51 uses the new robot 2d to cause the shared learning model 5 to execute the learned shared action, and acquires the robot's action data (step S41). In the initial learning process, action data was acquired for all the robots to be used, but here, only the action data for the new robot 2d is acquired.
[0096] The management unit 51 generates teacher data for learning the learning model of the action learning device 1 from the action data acquired in step S41 (step S42). This process is similar to step S22 of the initial learning process.
[0097] The management unit 51 uses the training data generated in step S42 to train the first learning model 3d and the second learning model 4d, which are part of the motion learning device 1 (step S43). The management unit 51 provides input training data to the first learning model 3d corresponding to the robot 2d and changes the parameters of the first learning model 3d and the second learning model 4d using the error between the output of the second learning model 4d corresponding to the robot 2d at that time and the output training data. In the initial training process, the hatched areas corresponding to the robots 2a to 2c shown in Figures 9B to 9D were trained. However, in the new robot addition process, the first learning models 3a to 3c, the second learning models 4a to 4c, and the shared learning model 5 for the existing robots 2a to 2c have already been trained. Therefore, when adding a new robot 2d, only the first learning model 3d and the second learning model 4d of the newly added robot 2d need to be trained using a series of shared motion sequences already learned in the shared learning model 5.
[0098] The management unit 51 stores the first learning model 3d and the second learning model 4d generated in step S43 in the robot 2d as behavior models, and stores the shared learning model 5 in the robot 2d as a behavior model (step S44). As a result, the robot 2d can use the shared learning model 5 configured with the stored behavior models to execute the shared learning behavior that has been learned by the other robots 2a to 2c in the configuration shown in Fig. 11C.
[0099] 11C is a diagram showing the configuration of a newly added robot 2d. At this time, sensing data from the external / internal world measurement unit 80 is input to the first learning model 3d, and an actuator 81 is driven based on predicted motion information output via the shared learning model 5 and the second learning model 4d. This makes it possible to reduce the learning load, and also makes it possible for the newly added robot 2d to execute high-quality actions that have already been acquired by the other robots 2a to 2c.
[0100] Individual robot learning processing The individual robot learning process of the action learning device 1 will be described below with reference to Fig. 12A to Fig. 12D. Here, tuning is performed on the robots 2a to 2c that have already learned in order to improve the sensing sensitivity, the quality and accuracy of the actions, the success rate, etc. Tuning here refers to minute parameter adjustments to the existing learning model, re-learning, additional learning, etc.
[0101] When the shared learning model 5 acquires new teaching data for a robot regarding a shared action that has already been learned, the management unit 51 uses the teaching data to perform either a first learning process in which only the first learning model and second learning model corresponding to this robot are trained, or a second learning process in which only the shared learning model 5 corresponding to the shared action is trained, or alternately performs the first learning and the second learning.
[0102] The processing of each part constituting the individual robot learning processing in Fig. 12A is the same as the processing explained so far, but the learning conditions are different. Here, the explanation will be given for robot 2a.
[0103] First, the management unit 51 uses the robot 2a to execute the shared action that has been learned in the shared learning model 5, and acquires the action data of the robot 2a (step S51). The method of acquiring the action data is the same as step S21 in Fig. 9A. Since the aim here is to improve performance, the management unit 51 collects action data that can be expected to improve quality, such as actions with a high success rate, actions that are robust to the environment, and actions that can be appropriately handled by sensing.
[0104] The management unit 51 may repeatedly collect motion data and select high-quality data from among them, or if high-quality motion is acquired by chance, the data may be used to perform tuning in step S21. Note that when the shared learning model 5 is updated by another robot and it is desired to update only the first learning model 3a and the second learning model 4a of the robot 2a accordingly, the management unit 51 may use existing motion data that has been used in the past.
[0105] The management unit 51 generates teacher data for training the action learning device 1 from the action data acquired in step S51 (step S52). This process is similar to step S22 of the initial learning process.
[0106] Next, in step S53, the process branches and moves to the following three steps depending on the update status.
[0107] First, consider the case where there is a shared learning model 5 that can generate high-quality movements with other robots. If the shared learning model 5 is updated by other robots and the quality of the movements of robot 2a is to be improved, in step S53, the management unit 51 selects a first learning model and a second learning model. Then, the management unit 51 trains the first learning model 3a and the second learning model 4a using the training data generated in step S52 (step S54), and the process returns to step S53.
[0108] If high-quality motion is acquired during the operation of the robot 2a, the management unit 51 selects a shared learning model in step S53. Then, the management unit 51 trains the shared learning model 5 using the training data generated in step S52 (step S55), and the process returns to step S53.
[0109] Here, the management unit 51 performs either the first learning in step S54, which trains the first learning model 3a and the second learning model 4a, or the second learning in step S55, which trains the shared learning model 5, or alternately performs the first learning and the second learning.
[0110] If the first learning model 3a, the shared learning model 5, and the second learning model 4a are trained simultaneously, the optimal structure for only the robot 2a will be learned. As a result, the structure in which common information is trained in the shared learning model 5 and the robot-specific input / output information processing is separated and trained in the first learning model 3 and the second learning model 4 will be destroyed.
[0111] Therefore, if high-quality movement data that cannot be achieved by other robots is obtained during step S51 and it is desired to train both the first learning model 3a, the second learning model 4, and the shared learning model 5, steps S54 and S55 must be executed alternately. After the desired behavior model is acquired in the robot 2a and the update process is completed, the shared learning model 5 and / or the first learning model 3a and the second learning model 4a are stored in the robot 2a as behavior models (step S56).
[0112] 12D is a diagram showing the configuration of robot 2a. At this time, sensing data from external / internal world measurement unit 80 is input to first learning model 3a, and actuator 81 is driven based on predicted motion information output via shared learning model 5 and second learning model 4a. This enables the robot 2a to perform high-quality movements using the newly saved shared learning model 5 and / or the movement models of the first learning model 3a and the second learning model 4a, in the configuration shown in Figure 12D, using the additional movement data obtained by the robot 2a.
[0113] In the present invention, the above method enables the sharing and transfer of learning information and learning models between multiple types of robots, particularly between robots with different mechanisms, structures, or characteristics. By sharing and transferring high-quality movements acquired by other robots to robots that have not yet learned those movements, the learning load, such as acquiring training data and tuning parameters, can be reduced, and high-quality movements can be easily acquired.
[0114] Specific examples of the effects of the present invention are described below. As an example, we will explain the application of the present invention to two types of robots: a smart robot with a high degree of freedom of movement and a wide variety of sensors, and a simple robot with the minimum configuration necessary to perform desired movements. Simple robots are lightweight and can be easily controlled using a remote device (even if they fail, they will not damage objects), and they require a small amount of movement data, making them easy to learn. However, they cannot acquire advanced movements. Smart robots are heavy, have complex and delicate systems, are difficult to control, and require a large amount of movement data, resulting in a heavy learning load. However, because smart robots have high sensing and movement capabilities, they can perform advanced, high-quality tasks that require skill. By using both robots and performing initial learning using simple movements once, the following advantages are achieved when learning new, specific desired movements.
[0115] First, the simple robot acquires motion data in step S31 of the unlearned motion learning process, and shares the motion with the smart robot in step S33. Next, the parameters of the shared learning model 5 acquired by the simple robot are used as initial values, and the smart robot is used to tune the shared learning model 5 in step S55 of the individual learning process. In this way, basic motions have been acquired by the simple robot, so the learning load can be reduced compared to when a large amount of motion data is acquired and learned from an initial state by the smart robot, and high-quality motions can be acquired more easily. Furthermore, if the high-quality motions acquired by the smart robot here are shared with the simple robot in step S34, it is easier to acquire motions of higher quality than motions obtained by the simple robot alone.
[0116] The configuration and effects of the present invention will be described below. [1] a plurality of first learning models (3a to 3c) that receive motion information at a certain time and convert it into motion feature quantities for each of a plurality of types of robots (2a to 2c), and that receive external world information at the same time and convert it into external world feature quantities; a shared learning model (5) that converts the motion feature quantity and external environment feature quantity output by each of the first learning models (3a to 3c) into a predicted motion feature quantity for the next time common to a plurality of types of the robots (2a to 2c); a plurality of second learning models (4a to 4c) that convert the predicted motion feature amount at the next time into predicted motion information for each of the plurality of types of robots (2a to 2c); a management unit (51) that uses teacher data related to the operation of each of the robots (2a to 2c) to train one of the first learning models (3a to 3c), the second learning models (4a to 4c), and the shared learning model (5) related to the robots (2a to 2c); A robot motion learning device (1) comprising:
[0117] This reduces the learning load on the learning model that controls multiple types of robots.
[0118] [2] the management unit (51) uses teacher data relating to movements not yet learned by each of the robots to train the shared learning model (5) while keeping the parameters of the first learning models (3a to 3c) and the second learning models (4a to 4c) fixed; 2. The robot action learning device (1) according to claim 1.
[0119] This reduces the learning load for new actions.
[0120] [3] the management unit (51) newly adds the first learning models (3a to 3c) and the second learning models (4a to 4c) corresponding to a robot to be newly added as a control target; 2. The robot action learning device (1) according to claim 1.
[0121] This reduces the learning load for each new robot movement.
[0122] [4] the management unit (51) uses teacher data of the robot to be added relating to the motion already learned by the shared learning model (5) to train the first learning models (3a to 3c) and the second learning models (4a to 4c) while keeping the parameters of the shared learning model (5) fixed; 2. The robot action learning device (1) according to claim 1.
[0123] This reduces the learning load on newly added robots.
[0124] [5] When the shared learning model (5) acquires new training data for the robot regarding the learned shared behavior, The management unit (51) uses the teacher data to perform either a first learning in which only the first learning model (3a to 3c) and the second learning model (4a to 4c) corresponding to the robot are learned, or a second learning in which only the shared learning model (5) corresponding to the shared action is learned, or alternately performs the first learning and the second learning. 2. The robot action learning device (1) according to claim 1.
[0125] This allows the shared learning model (5) to improve the accuracy of the learned shared actions.
[0126] [6] The shared learning model (5a, 5b) is provided for each of a plurality of shared actions, an action specifying unit (52) for selecting and executing one of the shared learning models (5a, 5b) provided for each of a plurality of shared actions; 6. The robot action learning device (1) according to claim 5, further comprising:
[0127] This allows for improved accuracy of each shared operation.
[0128] [7] the shared learning model (5) converts the external environment feature quantity output by each of the first learning models (3a to 3c) into a predicted external environment feature quantity for the next time that is common to the plurality of types of the robots (2a to 2c); 2. The robot action learning device (1) according to claim 1.
[0129] This makes it easy to determine whether learning by the learning model has converged.
[0130] [8] A robot motion learning device (1) according to claim 1; Multiple robots (2a-2c) and A robot action learning system (60) comprising:
[0131] This makes it possible to provide a system that reduces the learning load on a learning model that controls multiple types of robots.
[0132] [9] A robot motion learning method for learning motions of multiple types of robots (2a to 2c), a step in which a first learning model (3a to 3c) corresponding to a certain robot learns a process of converting motion information and external world information of the robot at a certain time into common motion feature quantities using teacher data related to the motion of the robot; a step in which a shared learning model (5) learns a time series relationship of common motion features related to motions common to the plurality of types of robots (2a to 2c) using the training data related to motions of the plurality of types of robots (2a to 2c); a step in which a second learning model (4a to 4c) corresponding to the robot uses the teacher data related to the robot's motion to learn a process of converting a predicted value of the common motion feature at a next time instant output by the shared learning model into predicted motion information of the robot at a next time instant; A robot motion learning method comprising:
[0133] This reduces the learning load on the learning model that controls multiple types of robots.
[0134] <<Variation>> The present invention is not limited to the above-described embodiments and includes various modifications. For example, the above-described embodiments have been described in detail to clearly explain the present invention, and are not necessarily limited to those including all of the described configurations. It is possible to replace part of the configuration of one embodiment with the configuration of another embodiment, and it is also possible to add the configuration of another embodiment to the configuration of one embodiment. Furthermore, it is also possible to add, delete, or replace part of the configuration of each embodiment with other configurations.
[0135] The above-described configurations, functions, processing units, processing means, etc. may be realized in part or in whole by hardware such as an integrated circuit. The above-described configurations, functions, etc. may be realized by software by a processor interpreting and executing a program that realizes each function. Information such as the programs, tables, and files that realize each function can be stored in a storage device such as a memory, a hard disk, or an SSD (Solid State Drive), or on a storage medium such as a flash memory card or a DVD (Digital Versatile Disk).
[0136] In each embodiment, the control lines and information lines shown are those that are considered necessary for the explanation, and not all control lines and information lines in the product are necessarily shown. In reality, it can be considered that almost all components are interconnected. [Explanation of symbols]
[0137] 1. Motion learning device 2a~2d Robot 3a~3d First learning model 4a~4d Second Learning Model 5. Shared Learning Model 51 Management Department 52 Operation specification section 61 Network 64 Robot motion teaching device 70 Processing unit 71 CPU 72 ROM 73 RAM 74 External Memory 78 System Bus 77 Communication Interface 75 Display section 76 Input section 80 External and Internal Measurement Unit 81 Actuator 91a, 91b, 91c External features 92a, 92b, 92c Internal features 99 Recursive Loops 93 External Features 94 Internal Features 95 Predicted external features 96 Predicted inner bound features
Claims
1. a plurality of first learning models that receive motion information at a certain time and convert it into motion feature quantities for each of a plurality of types of robots, and that receive external world information at the same time and convert it into external world feature quantities; a shared learning model that converts the motion feature quantity and the external environment feature quantity output by each of the first learning models into a predicted motion feature quantity for the next time point that is common to a plurality of types of the robots; a plurality of second learning models that convert the predicted motion feature amount at the next time into predicted motion information for each of the plurality of types of robots; a management unit that uses teacher data related to the operation of each of the robots to train one of the first learning model, the second learning model, and the shared learning model related to the robot; A robot motion learning device comprising:
2. the management unit trains the shared learning model using teacher data related to an operation that has not been learned by each of the robots, while keeping parameters of the first learning model and the second learning model fixed; 2. The robot action learning device according to claim 1.
3. the management unit newly adds the first learning model and the second learning model corresponding to a robot to be newly added as a control target; 2. The robot action learning device according to claim 1.
4. the management unit trains the first learning model and the second learning model using training data of the robot to be added regarding the motion already learned by the shared learning model, while keeping parameters of the shared learning model fixed; 2. The robot action learning device according to claim 1.
5. When the shared learning model acquires new robot training data related to the shared behavior it has already learned, The management unit uses the teacher data to perform either a first learning in which only a first learning model and a second learning model corresponding to the robot are learned, or a second learning in which only the shared learning model corresponding to the shared action is learned, or alternately performs the first learning and the second learning.
2. The robot action learning device according to claim 1.
6. the shared learning model is provided for each of a plurality of shared actions, an action designation unit that selects and executes one of the shared learning models provided for each of a plurality of shared actions; 6. The robot action learning device according to claim 5, further comprising:
7. the shared learning model converts the external environment feature quantity output by each of the first learning models into a predicted external environment feature quantity for the next time point common to the plurality of robots; 2. The robot action learning device according to claim 1.
8. The robot action learning device according to claim 1; Multiple types of robots, A robot motion learning system comprising:
9. A robot motion learning method for learning motions of a plurality of types of robots, comprising: a step in which a first learning model corresponding to a certain robot learns a process of converting motion information and external world information of the robot at a certain time into common motion features using training data related to motions of the certain robot; a step in which a shared learning model learns a time series relationship of common motion features related to motions common to the plurality of types of robots using the training data related to motions of the plurality of types of robots; a step in which a second learning model corresponding to the robot learns, using the teacher data related to the robot's motion, a process of converting a predicted value of the common motion feature at a next time instant output by the shared learning model into predicted motion information of the robot at a next time instant; A robot motion learning method comprising:
Citation Information
Patent Citations
General-purpose learned model generating method
JP2023089023A