Method for providing training data for a control unit of a robot arm and method for training a control unit of a robot arm

The method addresses the inefficiencies in robot arm training by generating robot arm movements from human demonstrations using contrastive learning and keyframes, ensuring accurate and efficient training through a model-based reinforcement learning algorithm.

WO2026002984A1PCT designated stage Publication Date: 2026-01-02SCHAEFFLER TECHNOLOGIES AG & CO KG +1
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
PCT/EP2025/067738
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2024-06-25
Filing Date
2025-06-24
Publication Date
2026-01-02

AI Technical Summary

Technical Problem

Existing methods for training robot arms through human demonstrations face challenges due to discrepancies in visual representation and the reliance on paired human-robot training data, leading to inefficiencies and errors in learning performance.

Method used

A method for providing training data to a robot arm controller using a model-based reinforcement learning algorithm, involving the generation of second video data from initial human arm movements through a transformation method, selection of keyframes, and applying contrastive learning with attention mechanisms to ensure accurate translation and reward-based training.

Benefits of technology

Enables efficient and reliable training of robot arm control units directly from human demonstrations with increased accuracy and reduced effort by utilizing selected keyframes and reward-based learning.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure EP2025067738_02012026_PF_FP_ABST
    Figure EP2025067738_02012026_PF_FP_ABST
Patent Text Reader

Abstract

The invention relates to a method for providing training data for a control unit of a robot arm, the control unit being designed to execute an algorithm for model-based reinforcement learning, the method comprising the following method steps: - providing first video data (10) representing a movement of a human arm; - generating second video data (12) by means of a transformation method depending on the first video data (10), the second video data (12) representing a movement of a robot arm, and the second video data (12) comprising a plurality of second individual frames; - selecting second individual frames from the second video data (12) as key frames (21); and - providing training data for the control unit comprising the second video data (12) and the selected key frames (21). The invention also relates to a method for training a control unit of a robot arm, which control unit is designed to execute an algorithm for model-based reinforcement learning (22), using training data obtained according to such a method for providing training data.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] SF2609P-WO-0013 P231437 DE1 Method for providing training data for a robot arm control unit and method for training a robot arm control unit. The invention relates to a method for providing training data for a robot arm control unit. A further object of the invention is a method for training a robot arm control unit. Background of the invention. In the following, the prior art relevant to the invention will first be outlined with reference to the sources listed below.[1] BD Argall, S. Chernova, M. Veloso, and B. Browning, “A survey of robot learning from demonstration,” Robotics and autonomous systems, vol. 57, no. 5, pp. 469–483, 2009.[2] J. Kober and J. Peters, “Policy search for motor primitives in robotics,” Advances in neural information processing systems, vol. 21, 2008.[3] T. Fitzgerald, E. Short, A. Goel, and A.Thomaz, “Human-guided trajectory adaptation fortool transfer,” in Proceedings of the 18th International Conference on Autonomous Agents andMultiAgent Systems, 2019, pp. 1350–1358.[4] M. Hazara and V. Kyrki, “Reinforcement learning for improving imitated in-contact skills,” in2016 IEEE-RAS 16th International Conference on Humanoid Robots (Humanoids). IEEE,2016, pp.194–201.[5] B. Akgun, M. Cakmak, J. W. Yoo, and A. L. Thomaz, “Trajectories and keyframes forkinesthetic teaching: A human-robot interaction perspective,” in Proceedings of the seventhannual ACM / IEEE international conference on Human-Robot Interaction, 2012, pp.391–398.[6] T. Zhang, Z. McCarthy, O. Jow, D. Lee, X. Chen, K. Goldberg, and P. Abbeel, “Deepimitation learning for complex manipulation tasks from virtual reality teleoperation,” in 2018IEEE International Conference on Robotics and Automation (ICRA). IEEE, 2018, pp. 5628–5635.[7] M. Sieb, Z. Xian, A. Huang, O. Kroemer, and K.Fragkiadaki, “Graphstructured visualimitation,” in Conference on Robot Learning. PMLR, 2020, pp.979–989.[8] X. B. Peng, A. Kanazawa, J. Malik, P. Abbeel, and S. Levine, “Sfv: Reinforcement learningof physical skills from videos,” ACM Transactions On Graphics (TOG), vol. 37, no. 6, pp. 1–14, 2018.SF2609P-WO-0013 P231437 DE2[9] X. B. Peng, E. Coumans, T. Zhang, T.-W. Lee, J. Tan, and S. Levine, “Learning agile roboticlocomotion skills by imitating animals,” arXiv preprint arXiv:2004.00784, 2020.

[0010] P. Sharma, D. Pathak, and A. Gupta, “Third-person visual imitation learning via decoupledhierarchical controller,” Advances in Neural Information Processing Systems, vol. 32, 2019.

[0011] S. Bahl, R. Mendonca, L. Chen, U. Jain, and D. Pathak, “Affordances from human videosas a versatile representation for robotics,” in Proceedings of the IEEE / CVF Conference onComputer Vision and Pattern Recognition, 2023, pp.13778–13790.

[0012] L. Smith, N. Dhawan, M. Zhang, P. Abbeel, and S.Levine, “Avid: Learning multi-stagetasks via pixel-level translation of human videos,” arXiv preprint arXiv:1912.04443, 2019.

[0013] K. Rao, C. Harris, A. Irpan, S. Levine, J. Ibarz, and M. Khansari, “Rl-cyclegan:Reinforcement learning aware simulation-to-real,” in Proceedings of the IEEE / CVF Conferenceon Computer Vision and Pattern Recognition, 2020, pp.11157–11166.

[0014] H. Xiong, Q. Li, Y.-C. Chen, H. Bharadhwaj, S. Sinha, and A. Garg, “Learning by watching:Physical imitation of manipulation skills from human videos,” in 2021 IEEE / RSJ InternationalConference on Intelligent Robots and Systems (IROS). IEEE, 2021, pp. 7827–7834.

[0015] J.-Y. Zhu, T. Park, P. Isola, and A. A. Efros, “Unpaired image-to-image translation usingcycle-consistent adversarial networks,” in Proceedings of the IEEE international conferenceon computer vision, 2017, pp.2223–2232.

[0016] A. Mandlekar, F. Ramos, B. Boots, S. Savarese, L. Fei-Fei, A. Garg, and D.Fox, “Iris: Implicit reinforcement without interaction at scale for learning control from offline robot manipulation data,” in 2020 IEEE International Conference on Robotics and Automation (ICRA). IEEE, 2020, pp. 4414–4420. The invention can be applied to the control of robot arms using learning from demonstration (LfD). Learning from demonstration can initialize an agent and accelerate the search for the optimal strategy [1]. It has been successfully demonstrated in various tasks [2]-[4]. Extensive work has been carried out on LfD, where human supervision is provided through kinesthetic instruction [5] and teleoperation [6]. With these two approaches, collecting demonstrations is laborious and slow. In the prior art, it is known to train robot arms by observing a human demonstrator.The challenge lies in the discrepancies between the visual representation of humans and robots. To address this problem, several recent studies have investigated this task using explicit pose and object recognition, with the key elements being trained by a pre-trained pose estimation module. The VEGs in [7] utilize 3D recognition to minimize the visual gap between the human and robot domains. SFV in [8] enables humanoid characters to learn skills from videos based on deep pose estimation, and the agent then learns to replicate the human reference movement. Similar to [8], the network in [9] also learns from animal videos and applies the observed movement to the simulation. Several other approaches using predictive modeling are also being investigated to minimize the human-robot domain gap

[0010] ,

[0011] .DHC

[0010] introduces a hierarchical structure in which a high-level module learns to generate sub-goals for the first person based on a video demonstration by a third party, and the low-level controller then predicts the actions required to achieve these sub-goals. VRB

[0011] trains a visual affordance model that can predict contact points and waypoints of the trajectory with the aid of future images to estimate where and how a person is likely to interact in the scene. However, the dependence on such paired human-robot training data limits the development of these methods. Furthermore, the indirect extraction of features from videos introduces errors that hinder learning performance in downstream tasks.Motivated by these limitations, many recent works in

[0012] -

[0014] employ generational models to transfer human demonstrations to the robot domain in an unpaired data environment. AVID

[0012] translates the human demonstrations into the robot domain. CycleGAN

[0015] then uses the generated video to construct a reward function for a model-based reinforcement learning algorithm. Similarly, RL-CyclyGAN

[0013] transforms a simulator image into a realistic image using the same CycleGAN framework

[0015] to bridge the gap between simulation and reality. LBW

[0014] translates each frame of a human video using UNIT

[0016] and performs unsupervised recognition of key points on the generated robot video.Keypoint-based trajectories can be used directly for calculating rewards and learning guidelines. However, RL-CyclyGAN

[0013] and LBW

[0014] aim to directly transfer demonstration videos to a simulated environment, but they still encounter the challenge of transferring from the simulation to the real environment. Disclosure of the Invention: Against this background, the task arises to enable reliable training of a robot arm's control device by means of demonstrations by a human with minimal effort.The problem is solved by a method for providing training data to a robot arm controller configured to execute a model-based reinforcement learning algorithm, comprising the following steps: ^Providing initial video data representing a movement of a human arm; ^Generating second video data using a transformation method dependent on the initial video data, wherein the second video data represents a movement of a robot arm and includes multiple second frames; ^Selecting second frames from the second video data as keyframes; and ^Providing training data to the controller comprising the second video data and the selected keyframes.The method according to the invention makes it possible to derive training data for training a robot arm's control unit directly from human demonstration videos. This training data comprises selected key frames that can significantly influence the training of the robot arm's control unit through model-based reinforcement learning. Due to the selected key frames, the efficiency of the training can be increased, thus enabling reliable training with reduced effort. According to an advantageous embodiment of the method, the process includes the following step preceding the provision of the first video data: Optically capturing a movement of a human arm, in particular by means of a video camera, to obtain the first video data.SF2609P-WO-0013 P231437 DE5 According to an advantageous embodiment of the method, the first video data comprises several first frames, and to generate the second video data, a corresponding second frame is generated for each first frame as part of the second video data. A transformation method based on a generative model can be used to generate the second video data, in which the movement of a human arm depicted in the first video data is translated into the corresponding movement of a robot arm. For example, convolutional layers and ResNet blocks can be used as part of the generator architecture in this transformation method. According to an advantageous embodiment of the method, a transformation method using a contrastive learning technique is used to generate the second video data.Contrastive learning is a machine learning method that focuses on extracting meaningful representations by comparing positive and negative pairs of instances. The goal is to use contrastive learning to enable the encoder-decoder structure to learn the relevant features of the input data—in this case, the first video data—and translate them into a meaningful representation—in this case, the second video data—that depicts the essential similarities and differences between the movements of a human arm and a robot arm. Therefore, by using contrastive learning, starting with the movement of a human arm depicted in the first video data, the generation of second video data containing the synthetically generated movement of a robot arm can be supported.According to an advantageous embodiment of the method, a attention module is used to detect critical features of the human arm's movement and determine the position of these features within the first video data to generate the second video data. The attention module comprises channel attention and spatial attention. Channel attention helps the system focus on the crucial features of the human movement, while spatial attention establishes the precise position of these features within the video images. These positions indicate the trajectory of the human arm, its relative positioning to objects, and the interaction space.By utilizing these two forms of attention, the system can improve the accuracy and precision of the transformation process, ensuring that the movements and actions of the robot arm in the generated second video data are consistent with the human demonstrations in the first video data. According to an advantageous embodiment of the method, lighting changes in the first video data are compensated for when generating the second video data. This compensation improves the quality of the conversion of human arm movements into robot arm movements. It also allows for greater robustness against environmental influences and background elements in the first video data.According to an advantageous embodiment of the method, a difference metric between two consecutive second frames is determined to select the second frame as the key frame from the second video data. This makes it possible to automate the selection of the key frames. Preferably, several second frames are selected as key frames. According to an advantageous embodiment of the method, the second frame is selected as the key frame if the difference metric of the second frame, relative to the frame preceding it, exceeds a predefined threshold.To solve the aforementioned problem, a method for training a robot arm's control unit is proposed. This unit is configured to execute a model-based reinforcement learning algorithm using training data obtained according to a previously described method for providing training data. The control unit training method thus provides a continuous, efficient process for reliably training the control unit through human demonstration. According to an advantageous embodiment of the method, the control unit receives a positive reward if the robot arm's movement resembles, or preferably matches, the selected key patterns.Conversely, the control unit receives a negative reward if the robot arm's movement deviates from the selected key images SF2609P-WO-0013 P231437 DE7. This ensures that the robot arm's movements, when trained by the control unit, mimic the movements of a human arm. Further details and advantages of the invention will be explained below with reference to the embodiment shown in the drawings. Figure 1 shows an embodiment of an inventive method for training in a block diagram. Figure 1 shows an embodiment of the inventive method for training a control unit of a robot arm, which is configured to execute a model-based reinforcement learning algorithm.The procedure is divided into two overarching processes: First, starting with initial video data representing the movement of a human arm, second video data is generated representing a corresponding movement of a robot arm (see Block 1). Then, the second video data is used to train the robot arm's control system with a model-based reinforcement learning algorithm (see Block 2). The procedure assumes that initial video data 10 representing the movement of a human arm is provided. For example, this initial video data 10 can be provided by optically capturing the movement of a human arm, particularly using a video camera. The initial video data 10 can contain movements of the human arm that correspond to the movements to be performed by the robot arm.Furthermore, the initial video data 10 can also contain other movements and non-movements, i.e., positions, of the human arm. It is also conceivable that the initial video data 10 represent movements of several different human arms. Based on and depending on this initial video data 10, a second video data 12 is then generated, representing a movement of a robot arm that corresponds to the movement of the human arm. This means, for example, that starting from a movement of a human arm being bent, as shown in the initial video data 10, a second video data 20 is generated, depicting a robot arm being bent in the same way. Furthermore, for example, a grasping movement of a human arm in the initial video data 10 is transformed into a grasping movement of the robot arm in the second video data 20.SF2609P-WO-0013 P231437 DE8 The first video data 10 consists of several consecutive first frames, and the second video data 12 consists of several consecutive second frames. A transformation method is used to generate the second video data 12, which utilizes a contrastive learning procedure 11. In this procedure, for each first frame of the first video data 10, a corresponding second frame of the second video data 12 is generated. Every human movement in the first video data 10 is translated into a movement of the robot arm in the second video data 12. The procedure 11 preferably uses several convolutional layers and ResNet blocks with an attention module as the generator. In addition, lighting changes in the first video data 10 are compensated for when generating the second video data.The second video data 12 forms the basis for the subsequent training of the robot arm's control system, see Block 2. For this purpose, several second frames from the second video data are first selected as so-called key frames 21. The selection of which of the second frames are to be considered key frames 21 is automated based on a difference metric. A second frame is selected as a key frame 21 if the difference metric of the second frame, relative to the frame preceding it, exceeds a predefined threshold. The second frames from the second video data selected as key frames 21 occupy a prominent position in the training. This is because the key frames serve as a reference for determining whether a positive or negative reward should be given within the framework of model-based reinforcement learning 22. If the position executed by the robot arm...If action 23 deviates from the position or action depicted in key image 21, a negative reward is issued. If the positions or actions 23 of the robot arm and key image 21 match, a positive reward is issued.

Claims

Method for providing training data to a control unit of a robot arm configured to execute a model-based reinforcement learning algorithm, comprising the following process steps: - Providing initial video data (10) representing a movement of a human arm; - Generating second video data (12) depending on the first video data (10), wherein the second video data (12) represent a movement of a robot arm, and wherein the second video data (12) comprise several second frames; Selecting second frames from the second video data (12) as keyframes (21); and Providing training data for the control device comprising the second video data (12) and the selected key frames (21), characterized in that a difference metric between two consecutive second frames is determined from the second video data (12) to select the second frame as the key frame (21).

2. Method according to claim 1, characterized by the following method step preceding the provision of the first video data (10): - Optical detection of a movement of a human arm, especially by means of a video camera, to obtain the first video data (10).

3. Method according to one of the preceding claims, characterized in that the first video data (10) comprise several first individual frames and, to generate the second video data (12), a corresponding second individual frame is generated for each first individual frame as part of the second video data (12).

4. Method according to one of the preceding claims, characterized in that a method of contrastive learning (11) is used to generate the second video data (12).

5. Method according to one of the preceding claims, characterized in that the second video data (12) is generated by means of an attention module. Critical features of the movement of the human arm are identified and a position of these critical features is determined in the first video data (10).

6. Method according to one of the preceding claims, characterized in that changes in lighting in the first video data (10) are compensated for in order to generate the second video data (12).

7. Method according to claim 1, characterized in that the second single image is selected as the key image (21) if the difference metric of the second single image, with respect to the single image preceding the second single image, exceeds a predetermined threshold value.

8. Method for training a control device of a robot arm, which is configured to execute a model-based reinforcement learning algorithm (22), with training data, characterized in that the training data are obtained according to a method according to one of the preceding claims.

9. Method according to claim 8, characterized in that a positive reward is specified to the control device if the movement of the robot arm resembles the selected key images (21).