Robot Trajectory Transfer Learning Across Simulation and Reality
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current methods for programming industrial robots are labor-intensive and require extensive training, leading to high costs and risks, especially for complex tasks, due to the gap between simulated and real-world environments during training, which necessitates time-consuming refinement of learned trajectories.
Innovation Solution
A system that uses a neural network-based transfer module to map simulated robot trajectories to real-world trajectories, allowing robots to imitate simulated tasks by comparing desired trajectories in both environments, thereby reducing the time required to learn new tasks and minimizing damage and operational costs.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Loss of time
If robot training is performed in simulation environment, then training cost and safety risk are reduced, but there is a gap between simulated and real-world trajectories requiring time-consuming refinement
Solution Approach 1:
The patent introduces a transfer module as an intermediary that learns to map simulated robot trajectories to real-world trajectories. This transfer module is trained on pairs of simulated and real trajectories, enabling the robot to directly transfer skills from simulation to reality without time-consuming fine-tuning, thus resolving the gap between simulation and real-world performance
Solution Approach 2:
The system transforms trajectories by learning parameter mappings between simulation and reality. The transfer module learns to transform simulated trajectory parameters (positions, orientations, velocities) into corresponding real-world parameters, allowing direct transfer of trained policies without retraining
2Reliability
If traditional robot programming is used, then task execution reliability is ensured, but extensive training is required leading to high costs and operational disruption
Solution Approach 1:
The patent uses simulated robot trajectories as copies or proxies for real-world trajectories. By training in simulation and using a transfer module to map these simulated trajectories to reality, the system avoids the need for extensive real-world programming and training, thereby maintaining reliability while improving productivity
Solution Approach 2:
The system performs preliminary training in a virtual simulation environment before deploying to the real robot. The transfer module is pre-trained on simulated trajectories, so when the real robot needs to perform a task, it can directly use the pre-learned mappings without requiring extensive on-site programming or trial-and-error learning
3Manufacturing precision
If robot learns tasks in real-world environment, then task performance accuracy is improved, but training time is very long and damage risk is high
Solution Approach 1:
The transfer module serves as an intermediary that bridges simulation and reality. It learns the mapping between simulated and real trajectories, allowing the robot to achieve real-world task performance accuracy by transferring knowledge from simulation without requiring long real-world training periods
Solution Approach 2:
The patent replaces the mechanical trial-and-error learning process in the real world with a computational approach. Instead of physically training the robot in the real world (which takes time and risks damage), the system uses computational models and transfer learning to achieve the same training objectives virtual
Data Source
Figure 1
Figure 2
Figure 3A
AI summary
A system for trajectories imitation for robotic manipulators is provided. The system includes an interface configured to receive a plurality of task descriptions, wherein the interface is configured to communicate with a real-world robot, a memory to store computer-executable programs including a robot simulator, a training module and a transfer module, and a processor, in connection with the memory. The processor is configured to perform training using the training module, for the task descriptions on the robot simulator, to produce a plurality of source policy with subgoals for the task descriptions. The processor performs training using the training module, for the task descriptions on the real-world robot, to produce a plurality of target policy with subgoals for the task descriptions, and update the parameters of the transfer module from corresponding trajectories with the subgoals for the robot simulator and real-world robot.