Robot Trajectory Transfer Between Simulation and Real-World Tasks
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current robotic training methods require extensive programming and training time, which can be costly and risky, especially for complex tasks, and there is a gap between simulated and real-world environments that necessitates time-consuming refinement when transferring learned tasks from simulation to real robots.
Innovation Solution
A system that compares and maps trajectories between simulated and real robots to determine correspondences, allowing real robots to imitate simulated tasks more efficiently, reducing training time and risk by using an interface, memory, and processor to perform training and update parameters for robotic manipulators.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Object-affected harmful factors
If robot training is performed in simulation environment, then training cost and safety risk are reduced, but there is a gap between simulation and real world requiring time-consuming fine-tuning
Solution Approach 1:
The patent creates a virtual copy of the real robot in a simulation environment that replicates its physical properties, dynamics, and sensorimotor characteristics. This virtual robot is trained to perform tasks, and the learned policies are then transferred to the real robot through systematic fine-tuning, eliminating the need to train the expensive real robot directly and reducing damage risk while minimizing adaptation time through accurate simulation-to-reality mapping
Solution Approach 2:
The patent performs preliminary training in the simulation environment before deploying to the real robot. The virtual robot learns task policies and skills in advance through safe, cost-free trials. This preliminary action in simulation prepares the robot with pre-trained capabilities that can be directly transferred to the real world, significantly reducing the fine-tuning time and avoiding damage to physical equipment during the learning process
2Manufacturing precision
If extensive programming is used for robot tasks, then task accuracy is improved, but programming complexity and time increase
Solution Approach 1:
The patent replaces traditional manual programming approaches with automated machine learning-based policy learning. Instead of manually writing complex control code to achieve precise task execution, the system uses reinforcement learning and imitation learning algorithms that automatically generate control policies through trial-and-error training in simulation. This substitution reduces programming complexity from hours of manual coding to automated training processes while maintaining or improving task accuracy through adaptive learning
3Reliability
If robot training is performed in real world, then task performance is improved, but training time and operating costs increase
Solution Approach 1:
The patent introduces a simulation environment as an intermediary between task specification and real-world execution. The simulation serves as a training ground where robots can learn policies without affecting real-world productivity. The learned policies are then transferred to real robots with minimal fine-tuning, achieving high task performance while maintaining real-world productivity because the bulk of training occurs in the virtual intermediary environment rather than stopping physical production
Data Source
AI summary
A system for trajectories imitation for robotic manipulators is provided. The system includes an interface configured to receive a plurality of task descriptions, wherein the interface is configured to communicate with a real-world robot, a memory to store computer-executable programs including a robot simulator, a training module and a transfer module, and a processor, in connection with the memory. The processor is configured to perform training using the training module, for the task descriptions on the robot simulator, to produce a plurality of source policy with subgoals for the task descriptions. The processor performs training using the training module, for the task descriptions on the real-world robot, to produce a plurality of target policy with subgoals for the task descriptions, and update the parameters of the transfer module from corresponding trajectories with the subgoals for the robot simulator and real-world robot.


