Adaptive Robot Control Blending RL With Feedback Signals

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Conventional feedback control techniques are brittle and inaccurate when dealing with complex physical interactions in industrial robotics, while reinforcement learning methods are data-inefficient and require extensive exploratory behavior, making them costly and time-consuming for practical industrial deployment.

Innovation Solution

The integration of adaptively weighted reinforcement learning and conventional feedback control, where control signals are compared for orthogonality and adjusted based on reward functions, and an iterative training approach is used to interleave simulated and real-world experiences to improve the control policy efficiently.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Adaptability or versatility

If reinforcement learning techniques are used to learn continuous robot controllers involving interactions with the physical environment, then the controller can handle complex physical interactions, but the training process becomes burdensome and time-consuming with substantial sample inefficiency

Engineering Contradiction:
Improveability to handle complex physical interactionsVSAvoidtraining time
Core Design Contradiction:
Adaptability or versatilityVSLoss of time

Solution Approach 1:

The patent applies preliminary action by pre-training the reinforcement learning controller in a simulated environment before deploying it to the physical robot. This allows the controller to learn complex physical interactions in advance without consuming real-world training time, resolving the contradiction between adaptability and training time loss.

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The patent creates a copy of the physical environment through simulation, allowing the reinforcement learning controller to be trained on a virtual replica rather than the actual physical system. This copying approach enables efficient training while maintaining the ability to handle complex physical interactions in the real world.

Inventive Principle:
Principle #26Copying

2Productivity

If reinforcement learning from scratch is used to train the control policy, then the controller can learn optimal strategies, but the process remains substantially data-inefficient and intractable

Engineering Contradiction:
Improvecontrol performanceVSAvoiddata efficiency
Core Design Contradiction:
ProductivityVSQuantity of substance

Solution Approach 1:

The patent performs preliminary action by pre-collecting demonstration data and pre-training the controller in simulation before actual deployment. This reduces the amount of real-world data needed during actual operation, improving data efficiency while maintaining high control performance.

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The patent introduces simulation environment and pre-collected demonstration data as intermediaries between the reinforcement learning algorithm and the physical robot. This intermediary approach allows the system to learn optimal strategies without requiring substantial real-world data, resolving the contradiction between productivity and data quantity requirements.

Inventive Principle:
Principle #24Intermediary (Mediator)

3Ease of manufacture

If manually tuned conventional controllers are used for deployment, then the controller can be deployed with existing techniques, but the process adds to costs and increases the time involved for robot deployment

Engineering Contradiction:
ImprovedeployabilityVSAvoiddeployment time
Core Design Contradiction:
Ease of manufactureVSLoss of time

Solution Approach 1:

The patent applies self-service by enabling the controller to automatically learn and adapt through reinforcement learning in simulation and real-world operation, eliminating the need for manual tuning by engineers. This self-service approach maintains ease of deployment while significantly reducing the time and cost associated with manual controller configuration.

Inventive Principle:
Principle #25Self-service

Data Source

PatentEP4017689B1Robotics control system and method for training said robotics control system
Publication Date: 2024.10.30 SIEMENS AG
  • EP4017689B1 patent drawingFigure 1
  • EP4017689B1 patent drawingFigure 2
  • EP4017689B1 patent drawingFigure 3

AI summary

Robotics control system (10) and method for training said robotics control system are provided. Disclosed embodiments make a gracefully blended utilization of Reinforcement Learning (RL) with conventional control by way of a dynamically adaptive interaction between respective control signals (20, 24) generated by a conventional feedback controller (18) and an RL controller (22). Additionally, disclosed embodiments make use of an iterative approach for training a control policy by effective use of virtual sensor and actuator data (60) interleaved with real-world sensor and actuator data (54). This is effective to reducing a training sample size to fulfill a blended control policy for the conventional feedback controller and the reinforcement learning controller. Disclosed embodiments may be used in a variety of industrial automation applications.