Robot Control Learning Using Simulation and Actual Data
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Collecting learning data for machine learning-based control modules for industrial robots is costly and risky, as it requires actual machine usage, leading to a gap between simulation and actual data, making it difficult to construct control modules operable in real environments.
Innovation Solution
A learning device that acquires and processes both simulation and actual data to train an extractor and controller separately, allowing the control module to convert sensor data into environmental information and derive control commands, bridging the gap between simulation and actual data.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Reliability
If learning data is collected using an actual machine, then the control module can be trained with real sensor data and state information, but cost and risk increase significantly
Solution Approach 1:
The patent creates a virtual copy of the actual machine through simulation, allowing learning data to be collected in a virtual environment that mirrors real-world conditions. This virtual machine copy enables data collection without physical risks or costs, while maintaining data quality suitable for training control modules
Solution Approach 2:
The patent introduces simulation technology as an intermediary between the actual machine and the learning process. This intermediary layer allows data to be generated indirectly through virtual representation, avoiding direct interaction with the actual machine during data collection phases
2Object-affected harmful factors
If learning data is collected using simulation, then cost and risk are reduced, but a gap between simulation data and actual data makes it difficult to construct operable control modules
Solution Approach 1:
The patent systematically varies simulation parameters to match real-world conditions, adjusting environmental factors, sensor characteristics, and machine states in the virtual environment to correspond with actual operating conditions. This parameter alignment reduces the gap between simulation and reality
Solution Approach 2:
The patent performs preliminary data processing and transformation on simulation data before using it for training. This includes preprocessing steps that align simulation data formats and characteristics with actual data expectations, preparing the virtual data in advance for effective control module training
3Extent of automation
If only simulation data is used for training, then data collection automation is achieved, but additional learning using actual machine data cannot be performed effectively
Solution Approach 1:
The patent segments the learning process into distinct phases: initial training using automated simulation data, followed by fine-tuning using actual machine data. This segmentation allows each phase to leverage its respective data source strengths while maintaining overall learning effectiveness
Solution Approach 2:
The patent designs a unified learning framework that can process both simulation data and actual machine data through the same control module architecture. This multi-functional approach allows the system to seamlessly transition between different data sources and perform multiple learning objectives
Data Source
AI summary
The disclosure is to constitute, while reducing a cost for collecting training data used in machine learning that makes a control module acquire an ability to control a robot device, the control module operatable in an actual environment by the machine learning. A learning device according to one aspect of the present invention executes machine learning of an extractor by using a first learning data set constituted by a combination of simulation data and first environmental information and a second learning data set constituted by a combination of actual data and second environmental information. Further, a learning device according to one aspect of the present invention executes machine learning of a controller by using a third learning data set constituted by a combination of third environmental information, state information, and a control command.


