Robot Skill Learning With Offline Pretraining for Precision Assembly

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Industrial robots face challenges in performing tight-tolerance assembly operations due to complex misalignments and uncertainties in part poses, leading to inefficient and potentially hazardous manual tuning of force control parameters, with existing reinforcement learning systems being slow and prone to failure.

Innovation Solution

A reinforcement learning controller is pre-trained offline using human demonstration data and then coupled to a compliance controller/robot system for self-learning, with co-training by a human operator to ensure successful assembly operations.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Reliability

If manual tuning of force control parameters is performed on real robotic systems, then the robot can be configured for assembly tasks, but the process is time-consuming and hazardous

Engineering Contradiction:
Improvesafety of tuning processVSAvoidtuning time
Core Design Contradiction:
ReliabilityVSLoss of time

Solution Approach 1:

The patent creates a virtual copy of the robotic system through simulation environment that replicates the physical system's dynamics and assembly task requirements. This virtual model allows all tuning operations to be performed safely without risking damage to actual robot or parts, while maintaining fidelity to the real system's behavior for accurate parameter transfer

Inventive Principle:
Principle #26Copying

Solution Approach 2:

The system performs force control parameter tuning in advance within the simulation environment before deploying to the real robotic system. By completing the time-consuming trial-and-error tuning process beforehand in a safe virtual setting, the actual robot deployment requires minimal or no tuning, significantly reducing on-site setup time and eliminating safety hazards

Inventive Principle:
Principle #10Preliminary action

2Extent of automation

If reinforcement learning systems are used for robot skill learning, then automation is improved, but training time is long and failure rate is high

Engineering Contradiction:
Improveautomation of skill learningVSAvoidtraining time
Core Design Contradiction:
Extent of automationVSLoss of time

Solution Approach 1:

The system pre-generates a comprehensive dataset of successful assembly demonstrations by having human operators perform tasks in the simulation environment. This pre-collected data serves as high-quality training material that captures correct assembly behaviors, allowing the reinforcement learning algorithm to learn from proven successful examples rather than discovering through numerous failed trials, thus dramatically reducing training time while maintaining full automation

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The system implements a feedback mechanism where the reinforcement learning controller receives reward signals based on its performance in the simulation environment. Successful assembly actions generate positive rewards that guide the learning process, enabling the system to efficiently identify effective strategies without extensive trial-and-error, thereby reducing training time while maintaining automation

Inventive Principle:
Principle #23Feedback

3Loss of time

If imitation learning systems are used for force control, then setup time is reduced, but robustness is poor and failure data overwhelms demonstration data

Engineering Contradiction:
Improvesetup timeVSAvoidlearning robustness
Core Design Contradiction:
Loss of timeVSReliability

Solution Approach 1:

The system introduces a simulation environment as an intermediary layer between human demonstration and the learning controller. This intermediate virtual environment allows demonstration data to be generated under controlled conditions with guaranteed successful outcomes, filtering out failure data before it reaches the learning system. The simulation acts as a mediator that translates human skills into clean training data, maintaining fast setup while improving robustness

Inventive Principle:
Principle #24Intermediary (Mediator)

Solution Approach 2:

The system creates multiple virtual copies of successful assembly demonstrations through simulation, generating abundant high-quality training data that reflects successful behaviors. By copying proven successful patterns from human operators in the virtual environment, the system ensures that demonstration data dominates the training set, preventing failure data from overwhelming the learning process while maintaining rapid setup through automated data generation

Inventive Principle:
Principle #26Copying

Data Source

PatentUS12384029B2Efficient method for robot skill learning
Publication Date: 2025.08.12 FANUC LTD
  • US12384029B2 patent drawing
  • US12384029B2 patent drawing
  • US12384029B2 patent drawing

AI summary

A method and system for robot skill learning for high precision assembly tasks employing a compliance controller. A reinforcement learning (RL) controller is pre-trained in an offline mode using human demonstration data, where several repetitions of the demonstration are performed while collecting state and action data for each repetition. The demonstration data is used to pre-train a neural network in the RL controller, with no interaction of the RL controller with the compliance controller/robot system. Following pre-training, the RL controller is moved to online production where it is coupled to the compliance controller/robot system in a self-learning mode. During self-learning, the neural network-based RL controller uses action, state and reward data to continue learning correlations between states and effective actions. Co-training is provided as needed during self-learning, where a human operator overrides the RL controller actions to ensure successful assembly operations, which improves the learned performance of the RL controller.