Robot Control With Haptic Reinforcement Learning for Contact-Rich Assembly

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current robotic systems lack adaptability and robustness in handling complex, unstructured manufacturing environments due to limitations in modeling system behaviors and generalizing control algorithms, especially in contact-rich tasks where traditional feedback control methods require explicit material models or lengthy trial-and-error processes.

Innovation Solution

The implementation of a reinforcement learning (RL) system using a mirror descent guided policy search (MDGPS) process that incorporates force/torque sensors for feedback, allowing robots to learn behaviors through interaction and generalize to new scenarios, combined with admittance force/torque control theory to enhance 'touch' and 'feel' capabilities during high-precision assembly tasks.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Reliability

If traditional feedback control methods are used, then control algorithms can be designed with explicit material models, but the system requires lengthy trial-and-error processes and lacks adaptability to new scenarios

Engineering Contradiction:
Improvecontrol accuracyVSAvoidtrial-and-error time
Core Design Contradiction:
ReliabilityVSLoss of time

Solution Approach 1:

The patent implements reinforcement learning with haptic feedback from force/torque sensors, where the robot continuously receives feedback about contact forces during assembly tasks and adjusts its control policy accordingly. This enables the system to learn from interaction rather than relying on pre-programmed trial-and-error processes, significantly reducing training time while maintaining control accuracy.

Inventive Principle:
Principle #23Feedback

Solution Approach 2:

The robot autonomously learns assembly skills through self-directed interaction with the environment using reinforcement learning. The system independently optimizes its control policy by processing haptic feedback and updating its neural network parameters, eliminating the need for extensive manual programming or supervised trial-and-error procedures.

Inventive Principle:
Principle #25Self-service

2Adaptability or versatility

If reinforcement learning with haptic feedback is implemented, then robots can autonomously acquire skills and adapt to new environments, but the system complexity increases due to deep neural networks and continuous learning processes

Engineering Contradiction:
Improveadaptability to new environmentsVSAvoidsystem complexity
Core Design Contradiction:
Adaptability or versatilityVSDevice complexity

Solution Approach 1:

The patent introduces haptic feedback from force/torque sensors as an intermediary that bridges the robot's actions and environmental responses. This tactile information serves as a rich feedback signal that enables the reinforcement learning system to learn complex assembly skills without requiring overly complicated system architectures, as the haptic feedback encapsulates essential interaction dynamics.

Inventive Principle:
Principle #24Intermediary (Mediator)

Solution Approach 2:

The patent replaces traditional mechanical control systems with reinforcement learning-based control. Instead of using complex mechanical mechanisms and explicit material models, the system uses deep neural networks that learn optimal control policies directly from haptic feedback, simplifying the overall system architecture while enhancing adaptability.

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

3Manufacturing precision

If force/torque sensors are used for haptic feedback, then robots can improve 'touch' and 'feel' capabilities during assembly, but the manufacturing cost increases

Engineering Contradiction:
Improveassembly precisionVSAvoidmanufacturing cost
Core Design Contradiction:
Manufacturing precisionVSEase of manufacture

Solution Approach 1:

The patent leverages force/torque sensors that serve multiple functions: they provide haptic feedback for reinforcement learning, enable precise contact force measurement during assembly, and offer tactile information for skill acquisition. This multi-functionality justifies the sensor cost by eliminating the need for separate measurement systems and enhancing assembly precision across multiple task types.

Inventive Principle:
Principle #6Universality (Multi-functionality)

Data Source

PatentEP3743250B1Reinforcement learning for contact-rich tasks in automation systems
Publication Date: 2024.12.18 SIEMENS AG
  • EP3743250B1 patent drawingFigure 1A~1B
  • EP3743250B1 patent drawingFigure 2A~2B
  • EP3743250B1 patent drawingFigure 3

AI summary

Systems and methods for controlling robots including industrial robots. A method includes executing (402) a program (550) to control a robot (102) by the robot control system (120, 500). The method includes receiving (404) robot state information (554). The method includes receiving (406) force torque feedback (556) inputs from a sensor (554) on the robot (102). The method includes producing (410) a robot control command for the robot (102) based on the robot state information (554) and the force torque feedback (556) inputs. The method includes controlling (412) the robot (102) using the robot control command.