Action learning system
Patent Information
- Application Number
- CN202610261742.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2025-03-28
- Filing Date
- 2026-03-05
- Publication Date
- 2026-09-29
AI Technical Summary
根据该方式的动作学习系统,能够容易使人型机器人再现人类作业者的动作。
Smart Images

Figure CN122829807A_ABST
Abstract
Description
Technical Field
[0001] This disclosure relates to motion learning systems. Background Technology
[0002] As a method for enabling humanoid robots to reproduce human actions, methods using hidden Markov models are known (e.g., Japanese Patent Application Publication No. 2004-330361). Summary of the Invention
[0003] However, using Hidden Markov Models requires computing state transition sequences, making the generation of motion data for humanoid robots to mimic human workers time-consuming. Therefore, there is a need for a technique that can easily perform learning for humanoid robots to imitate human workers' actions.
[0004] This disclosure can be implemented in the following ways.
[0005] (1) According to one aspect of the present disclosure, a motion learning system is provided. The motion learning system comprises: a humanoid robot; a camera equipped on the humanoid robot and capturing robot-view images, the robot-view images being images from a first-person perspective of the humanoid robot; and a learning unit that learns the motions of the humanoid robot based on operator-view images and the robot-view images to reduce the difference between the motions of a demonstration operator and the motions of the humanoid robot, the operator-view images being images from a first-person perspective of the demonstration operator captured during the period when the demonstration operator, as a human operator, performs a predetermined motion. According to this motion learning system, it is easy to learn how to make humanoid robots imitate the actions of human workers. (2) In the above-described motion learning system, it may include a field of view score calculation unit and an action score calculation unit. The field of view score calculation unit calculates a field of view score as the similarity between the operator's view image and the robot's view image. The action score calculation unit uses the robot's view image to estimate the robot action as the work content of the humanoid robot and calculates an action score as the similarity between the operator's action and the robot action. The operator's action is the work content of the demonstration operator estimated using the operator's view image. The learning unit learns the humanoid robot's action to improve the field of view score and the action score. According to this motion learning system, it is easy to learn how to make humanoid robots imitate the movements of human workers with high precision. (3) In the above-described motion learning system, the field of view score calculation unit may calculate the field of view score by comparing the objects presented in the operator's view image with the objects presented in the robot's view image. According to this motion learning system, the visual field score calculation unit can calculate the visual field score more accurately. (4) In the above-mentioned action learning system, the learning unit may perform reinforcement learning of the robot action model, wherein the robot action model is a learning model that outputs action control information for controlling the action of the humanoid robot when the robot's perspective image is input. (5) In the above-described motion learning system, a motion control unit may be provided, which controls the motion of the humanoid robot based on the learning results of the learning unit. Based on this motion learning system, humanoid robots can easily reproduce the actions of human workers. Attached Figure Description
[0006] The features, advantages, and technical and industrial significance of exemplary embodiments of the present invention are described below with reference to the accompanying drawings, in which the same reference numerals denote the same elements, wherein: Figure 1 This is an explanatory diagram showing the general structure of a motion learning system; Figure 2 This is an explanatory diagram showing the general structure of the control device; and Figure 3 This is a flowchart illustrating a learning method used to enable a humanoid robot to mimic the actions of a demonstrator operator. Detailed Implementation A. First implementation method:
[0007] Figure 1 This is an explanatory diagram showing the schematic structure of the motion learning system 10. The motion learning system 10 includes: a humanoid robot 100, a camera 200 mounted on the humanoid robot 100, and a control device 300. The motion learning system 10 is a system for learning to enable the humanoid robot 100 to imitate the actions of a human worker. The humanoid robot 100 is also referred to as a humanoid robot.
[0008] Camera 200 captures images of the humanoid robot 100 from a first-person perspective. In this embodiment, camera 200 is a wearable camera mounted on the head of the humanoid robot 100. In this embodiment, camera 200 is a so-called FPV (First-Person View) camera, which is detachably mounted on the humanoid robot 100. Alternatively, camera 200 may be integrally mounted on the head of the humanoid robot 100. Camera 200 captures images of the surroundings of the humanoid robot 100 from the perspective of the humanoid robot 100. Hereinafter, the images captured by camera 200 will be referred to as robot perspective images.
[0009] Figure 2 This is an explanatory diagram showing a schematic structure of the control device 300. The control device 300 is a computer including a processor 301, a memory 302, an input / output interface 303, and an internal bus 304. The processor 301, memory 302, and input / output interface 303 are connected bidirectionally via the internal bus 304. A camera 200 is connected to the input / output interface 303. The camera 200 can be connected to the input / output interface 303 via wired communication or wireless communication. In this embodiment, the control device 300 is mounted on a humanoid robot 100.
[0010] In addition to non-volatile memories such as EEPROM (Electrically Erasable Programmable Read-Only Memory), memory 302 also includes RAM (Random Access Memory). Memory 302 stores operator-view images 321, motion models 322, robot motion models 323, and control programs (not shown).
[0011] The worker's perspective image 321 is a first-person view of the demonstrator worker, captured during a demonstration worker's performance of pre-defined actions. The worker's perspective image 321 is captured by a wearable camera mounted on the demonstrator worker's head. This wearable camera is an FPV camera, which captures images of the demonstrator worker's surroundings from their perspective. Preferably, the position of the wearable camera mounted on the demonstrator worker is the same as the position of the camera 200 on the humanoid robot 100. For example, if the wearable camera is mounted on the left side of the demonstrator worker's head, the camera 200 is preferably mounted on the left side of the humanoid robot 100's head. The pre-defined actions performed by the demonstrator worker are actions related to the manufacture of an item. These actions include, in addition to actions involving the item itself such as inserting connectors, tightening bolts, holding items by hand, and moving parts, actions accompanying the manufacture of the item such as walking movements performed by the demonstrator worker. The memory 302 stores one worker's perspective image 321. In addition, the pre-defined actions performed by the demonstrator can be actions other than those involved in the manufacture of the item.
[0012] Action model 322 is a learning model that takes worker-view image 321 or robot-view image as input and outputs the tasks performed by the worker or humanoid robot 100 during the period when the input image is captured. Action model 322 is a learning model that has been pre-learned using a learning dataset through machine learning. As action model 322, for example, a convolutional neural network learned using a teacher-trained learning dataset can be used. The learning dataset, for example, includes worker-view image and labels representing the tasks performed by the worker during the period when the worker-view image is captured.
[0013] The robot motion model 323 is a learning model that outputs motion control information for controlling the movements of the humanoid robot 100 when inputting robot-view images. The motion control information includes, for example, command values such as the angles of the joints of the humanoid robot 100's arms and legs, the direction of movement of the arms and legs, and the speed of movement of the arms and legs.
[0014] The processor 301 executes the control program stored in the memory 302. Thus, the processor 301 functions as the operator action estimation unit 311, the field of vision score calculation unit 312, the action score calculation unit 313, the learning unit 314, and the motion control unit 315.
[0015] The worker action estimation unit 311 uses the worker's perspective image 321 to estimate the worker's actions as a demonstration of the worker's work. Examples of worker actions include inserting a connector, tightening a bolt, and holding an object by hand. The worker action estimation unit 311 uses a motion model 322 to estimate the worker's actions. Specifically, the worker action estimation unit 311 inputs the worker's perspective image 321 into the motion model 322 and estimates the work content output from the motion model 322 as the worker's actions performed by the demonstration worker during the period when the worker's perspective image 321 was captured.
[0016] The field-of-view score calculation unit 312 calculates a field-of-view score, which represents the similarity between the operator's view image 321 and the robot's view image, by comparing the operator's view image 321 with the robot's view image. Specifically, the field-of-view score calculation unit 312 calculates the field-of-view score by comparing objects presented in the operator's view image 321 with objects presented in the robot's view image. For example, the field-of-view score calculation unit 312 compares the positions of the demonstrator's hand, the connector held by the demonstrator, and the work machine presented in the operator's view image 321 with the positions of the humanoid robot 100's hand, the connector held by the humanoid robot 100, and the work machine presented in the robot's view image. Thus, the field-of-view score calculation unit 312 calculates the field-of-view score. The higher the field-of-view score, the more similar the operator's view image 321 is to the robot's view image.
[0017] The action score calculation unit 313 uses the robot's perspective image to estimate the robot actions that constitute the work content of the humanoid robot 100. Examples of robot actions include inserting a connector, tightening a bolt, and holding an object by hand. The action score calculation unit 313 inputs the robot's perspective image into the action model 322 and estimates the robot actions performed by the humanoid robot 100 during the period when the robot's perspective image was captured, based on the work content output from the action model 322. Furthermore, the action score calculation unit 313 calculates an action score as a measure of the similarity between the operator's actions and the robot's actions. The action score calculation unit 313 calculates the action score by determining whether the operator's actions estimated by the operator's action estimation unit 311 are consistent with the robot actions estimated by the action score calculation unit 313. The higher the action score, the more similar the operator's actions and robot actions are. For example, the action score calculation unit 313 sets the action score to 1 when the operator's actions and robot actions are consistent, and sets the action score to 0 when the operator's actions and robot actions are inconsistent.
[0018] The learning unit 314 learns the movements of the humanoid robot 100 based on the operator's perspective image 321 and the robot's perspective image, thereby reducing the difference between the movements of the demonstration operator and the movements of the humanoid robot 100. In this embodiment, the learning unit 314 learns the movements of the humanoid robot 100 to improve the field of view score and action score. The learning unit 314 learns the movements of the humanoid robot 100 by performing reinforcement learning on the robot motion model 323. The learning unit 314 performs machine learning of the robot motion model 323, for example, through reinforcement learning such as Q-learning or Monte Carlo methods. Alternatively, the learning unit 314 can also perform machine learning of the robot motion model 323 by using reinforcement learning such as Deep Q-Network (DQN) of deep learning.
[0019] The motion control unit 315 controls the movements of the humanoid robot 100 based on the learning results of the learning unit 314. Specifically, the motion control unit 315 controls the movements of the humanoid robot 100 based on the robot motion model 323 learned by the learning unit 314. The motion control unit 315 controls the movements of each part of the humanoid robot 100 according to the motion control information output from the robot motion model 323.
[0020] Figure 3 This is a flowchart illustrating a learning method for enabling a humanoid robot 100 to mimic the actions of a demonstration worker. In S10, the worker action estimation unit 311 uses the worker's perspective image 321 to estimate the worker actions performed by the demonstration worker during the period when the worker's perspective image 321 was captured. The worker action estimation unit 311 inputs the worker's perspective image 321 into the action model 322 and estimates the worker actions performed by the demonstration worker during the period when the worker's perspective image 321 was captured from the work content output by the action model 322.
[0021] In S20, the motion control unit 315 controls the movements of the humanoid robot 100 based on the robot motion model 323. The camera 200 captures robot-view images of the humanoid robot 100 during its actions. These captured robot-view images are sent to the control device 300 in real time. The motion control unit 315 inputs the robot-view images captured by the camera 200 into the robot motion model 323 and controls the movements of the humanoid robot 100 according to the motion control information output from the robot motion model 323. The robot-view images sent from the camera 200 are stored in the memory 302.
[0022] In S30, the field-of-view score calculation unit 312 calculates the field-of-view score by comparing the operator's view image 321 with the robot's view image captured in S20. The field-of-view score calculation unit 312 compares the operator's view image 321 with the robot's view image at multiple moments where the elapsed time from when the operator begins a pre-defined action is equal to the elapsed time from when the humanoid robot 100 begins its action. The field-of-view score calculation unit 312 compares the operator's view image 321 with the robot's view image at pre-defined time intervals. For example, the field-of-view score calculation unit 312 compares the operator's view image 321 with the robot's view image captured by the camera 200 in real time. The calculated field-of-view score is stored in the memory 302. Alternatively, the field-of-view score calculation unit 312 can also calculate the field-of-view score by comparing the operator's view image 321 with the robot's view image captured by the camera 200 in real time.
[0023] In S40, the action score calculation unit 313 estimates the robot actions performed by the humanoid robot 100 in S20. The action score calculation unit 313 inputs the robot perspective images captured in S20 into the action model 322, and estimates the work content output by the action model 322 as the robot actions performed by the humanoid robot 100 in S20.
[0024] In S50, the action score calculation unit 313 calculates an action score representing the similarity between the operator's action estimated in S10 and the robot's action estimated in S40. The calculated action score is stored in the memory 302.
[0025] In S60, the learning unit 314 performs reinforcement learning on the robot motion model 323 to reduce the difference between the actions of the demonstrator and the actions of the humanoid robot 100. The learning unit 314 uses field of view scores and action scores as rewards in the reinforcement learning, and performs reinforcement learning on the robot motion model 323 by having the humanoid robot 100 perform actions that yield higher rewards.
[0026] In S70, the control device 300 determines whether the learning of the robot motion model 323 has been completed. For example, if the field of view score calculated in S30 is higher than a predetermined reference value and the action score calculated in S50 is 1, the control device 300 determines that the learning of the robot motion model 323 has been completed. If the learning of the robot motion model 323 is determined to be completed, S80 is executed. If the learning of the robot motion model 323 is determined not to be completed, the process returns to S20.
[0027] In S80, the motion control unit 315 controls the movements of the humanoid robot 100 based on the learning results of the learning unit 314. The motion control unit 315 controls the movements of the humanoid robot 100 based on the learned robot motion model 323, enabling the humanoid robot 100 to perform actions related to the manufacture of the item. Specifically, the motion control unit 315 inputs the robot's perspective images captured by the camera 200 into the learned robot motion model 323 in real time. The motion control unit 315 controls the movements of the humanoid robot 100 according to the motion control information output from the robot motion model 323.
[0028] The motion learning system 10 described above in the first embodiment includes a humanoid robot 100, a camera 200 for capturing images from the robot's perspective, and a learning unit 314. The learning unit 314 learns the motions of the humanoid robot 100 based on the operator's perspective image 321 and the robot's perspective image, thereby reducing the difference between the actions of a demonstration operator and the actions of the humanoid robot 100. Therefore, it is easy to perform learning for the humanoid robot 100 to imitate the actions of a human operator.
[0029] In this embodiment, the field-of-view score calculation unit 312 calculates a field-of-view score, which is the similarity between the operator's view image 321 and the robot's view image. The action score calculation unit 313 uses the robot's view image to estimate the robot's actions and calculates an action score, which is the similarity between the operator's actions and the robot's actions. The learning unit 314 learns the actions of the humanoid robot 100 to increase the field-of-view score and the action score. Therefore, it is easy to perform learning to enable the humanoid robot 100 to imitate the actions of a human operator with high precision.
[0030] Furthermore, in this embodiment, the field-of-view score calculation unit 312 calculates the field-of-view score by comparing the objects presented in the operator's view image 321 with the objects presented in the robot's view image. Therefore, the field-of-view score calculation unit 312 can calculate the field-of-view score more accurately.
[0031] Furthermore, in this embodiment, the motion control unit 315 controls the movements of the humanoid robot 100 based on the learning results of the learning unit 314. Therefore, the humanoid robot 100 can easily reproduce the movements of a human worker. B. Other implementation methods:
[0032] (B-1) In the above embodiment, the worker's viewpoint image 321 is pre-stored in the memory 302. Alternatively, the worker's viewpoint image 321 may not be pre-stored in the memory 302. In this case, the control device 300 acquires the worker's viewpoint image 321 from an external storage device or the like. Alternatively, the control device 300 may directly acquire the worker's viewpoint image 321 from a wearable camera installed on the demonstration worker.
[0033] (B-2) In the above embodiment, the motion learning system 10 includes a worker action estimation unit 311. In contrast, the motion learning system 10 may also lack a worker action estimation unit 311. In this case, the worker's actions are pre-stored in the memory 302.
[0034] (B-3) In the above embodiment, the field of view score calculation unit 312 calculates the field of view score by comparing the objects presented in the operator's view image 321 with the objects presented in the robot's view image. Alternatively, the field of view score calculation unit 312 can also calculate the field of view score by comparing the direction the demonstrator is facing with the direction the humanoid robot 100 is facing. The direction the demonstrator is facing is estimated using the operator's view image 321, and the direction the humanoid robot 100 is facing is estimated using the robot's view image.
[0035] (B-4) In the above embodiment, the action learning system 10 includes a field-of-view score calculation unit 312. Conversely, the action learning system 10 may also lack a field-of-view score calculation unit 312. In this case, the learning unit 314 learns the actions of the humanoid robot 100 to increase the action score. Specifically, the learning unit 314 uses the action score as a reward in reinforcement learning, performing reinforcement learning on the robot action model 323 by having the humanoid robot 100 perform actions that yield higher rewards.
[0036] (B-5) In the above embodiment, the motion learning system 10 includes an action score calculation unit 313. Conversely, the motion learning system 10 may also lack an action score calculation unit 313. In this case, the learning unit 314 learns the actions of the humanoid robot 100 to increase the field of view score. Specifically, the learning unit 314 uses the field of view score as a reward in reinforcement learning, and performs reinforcement learning of the robot motion model 323 by having the humanoid robot 100 perform actions that yield higher rewards.
[0037] (B-6) In the above embodiment, the motion control unit 315 controls the motion of the humanoid robot 100 based on the learning results of the learning unit 314. Conversely, the motion control unit 315 may also control the motion of the humanoid robot 100 without relying on the learning results of the learning unit 314. That is, Figure 3S80 of the learning method shown can also be omitted.
[0038] (B-7) In the above embodiment, the control device 300 is mounted on the humanoid robot 100. Alternatively, the control device 300 may be located outside the humanoid robot 100 and connected to the humanoid robot 100 via wired or wireless communication.
[0039] This disclosure is not limited to the embodiments described above, and can be implemented in various structures without departing from its spirit. For example, technical features in embodiments corresponding to the technical features in the various embodiments described in the Summary of the Invention section can be appropriately replaced or combined to solve some or all of the above-mentioned problems. Or, to achieve some or all of the above-mentioned effects, they can be appropriately replaced or combined. In addition, if a technical feature is not described as an essential feature in this specification, it can be appropriately deleted.
Claims
1. A motion learning system, comprising: Humanoid robots; A camera, equipped on the humanoid robot, captures images from the robot's perspective, which are first-person images of the humanoid robot; and The learning unit learns the humanoid robot's movements based on the operator's perspective images and the robot's perspective images to reduce the difference between the demonstration operator's movements and the humanoid robot's movements. The operator's perspective images are first-person images taken by the demonstration operator, who is acting as a human operator, during the performance of a pre-defined movement.
2. The action learning system according to claim 1, wherein, The motion learning system includes a vision score calculation unit and an action score calculation unit. The field of view score calculation unit calculates a field of view score as the similarity between the operator's view image and the robot's view image. The action score calculation unit uses the robot's perspective imagery to estimate robot actions that constitute the humanoid robot's work content, and calculates an action score as the similarity between the operator's actions and the robot actions. The operator's actions are the work content of the demonstration operator estimated using the operator's perspective imagery. The learning unit learns the actions of the humanoid robot to improve the field of vision score and the action score.
3. The action learning system according to claim 2, wherein, The field of view score calculation unit calculates the field of view score by comparing the objects presented in the operator's view image with the objects presented in the robot's view image.
4. The action learning system according to claim 1, wherein, The learning unit performs reinforcement learning on the robot motion model, which is a learning model that outputs motion control information for controlling the humanoid robot's movements when the robot's perspective image is input.
5. The action learning system according to claim 1, wherein, The motion learning system includes a motion control unit, which controls the motion of the humanoid robot based on the learning results of the learning unit.
Citation Information
Patent Citations
Method for recognizing / generating motion data by hidden markov model, and motion controlling method using the same and its controlling system
JP2004330361A