Simulated Policy Learning for Robot Parameter Tuning
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Tuning parameters of industrial application programs, such as robot application programs, is tedious and time-consuming, often delaying production, and existing machine learning methods struggle with insufficient data and the reality gap between simulations and real applications, leading to inefficient and potentially dangerous learning processes.
Innovation Solution
A method and system that utilizes simulated applications to generate and refine candidate policies through machine learning, collecting performance data to optimize parameters for real applications, reducing the need for physical robots and time, and minimizing the reality gap.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Productivity
If machine learning is applied to learn basic motion in robot applications, then the task to be solved is improved, but the amount of quality data required is insufficient
Solution Approach 1:
The patent creates a virtual copy of the real robot application through simulation. The virtual robot application replicates the physical environment, robot mechanics, and task requirements, allowing machine learning to occur in the virtual domain where data can be generated without physical constraints. This copying approach enables abundant data generation while avoiding the data scarcity problem in physical systems.
Solution Approach 2:
The patent performs preliminary machine learning and data generation in the virtual environment before deploying to the real system. By conducting explorative learning and parameter optimization in advance within the simulated application, the system prepares optimized parameters that can be directly transferred to the real robot, eliminating the need for extensive on-site data collection and tuning.
2Adaptability or versatility
If explorative learning is used to adapt to environmental change, then adaptability is improved, but the learning process becomes expensive and dangerous
Solution Approach 1:
The patent introduces a virtual robot application as an intermediary between the real robot system and the machine learning process. This intermediary environment allows explorative learning to occur safely, where dangerous or costly trial-and-error experiments can be conducted without affecting physical hardware. The virtual environment serves as a buffer that protects the real system while enabling comprehensive adaptability testing.
3Loss of time
If simulation based machine learning is used, then safety and speed are improved, but the reality gap between simulation and real application increases
Solution Approach 1:
The patent designs the virtual robot application to be universally representative of multiple real-world scenarios. By creating a simulation environment that captures essential physics, mechanics, and task dynamics applicable across different conditions, the learned parameters achieve broad transferability. The virtual environment is constructed to maintain universal validity while enabling rapid experimentation.
Solution Approach 2:
The patent implements feedback mechanisms where performance results from the virtual application are used to iteratively refine the machine learning model and parameter optimization. This feedback loop ensures that the virtual learning process remains aligned with real-world objectives, continuously improving transferability while maintaining the speed advantages of simulation-based training.
Data Source
AI summary
A method for applying machine learning to an application includes: a) generating a candidate policy by a learner; b) executing a program in at least one simulated application based on a set of candidate parameters provided based on the candidate policy and a state of the at least one simulated application, execution of the program providing interim results of tested sets of candidate parameters based on a measured performance information of the execution of the program; c) collecting a predetermined number of interim results and providing an end result based on a combination of the candidate parameters and/or the state with the measured performances information by a trainer; and d) generating a new candidate policy by the learner based on the end result.


