Adaptive Robot Control Blending RL With Feedback Signals
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Conventional feedback control techniques are brittle and inaccurate when dealing with complex physical interactions in industrial robotics, while reinforcement learning methods are data-inefficient and require extensive exploratory behavior, making them costly and time-consuming for practical industrial deployment.
Innovation Solution
The integration of adaptively weighted reinforcement learning and conventional feedback control, where control signals are compared for orthogonality and adjusted based on reward functions, and an iterative training approach is used to interleave simulated and real-world experiences to improve the control policy efficiently.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Adaptability or versatility
If reinforcement learning techniques are used to learn continuous robot controllers involving interactions with the physical environment, then the controller can handle complex physical interactions, but the training process becomes burdensome and time-consuming with substantial sample inefficiency
Solution Approach 1:
The patent applies preliminary action by pre-training the reinforcement learning controller in a simulated environment before deploying it to the physical robot. This allows the controller to learn complex physical interactions in advance without consuming real-world training time, resolving the contradiction between adaptability and training time loss.
Solution Approach 2:
The patent creates a copy of the physical environment through simulation, allowing the reinforcement learning controller to be trained on a virtual replica rather than the actual physical system. This copying approach enables efficient training while maintaining the ability to handle complex physical interactions in the real world.
2Productivity
If reinforcement learning from scratch is used to train the control policy, then the controller can learn optimal strategies, but the process remains substantially data-inefficient and intractable
Solution Approach 1:
The patent performs preliminary action by pre-collecting demonstration data and pre-training the controller in simulation before actual deployment. This reduces the amount of real-world data needed during actual operation, improving data efficiency while maintaining high control performance.
Solution Approach 2:
The patent introduces simulation environment and pre-collected demonstration data as intermediaries between the reinforcement learning algorithm and the physical robot. This intermediary approach allows the system to learn optimal strategies without requiring substantial real-world data, resolving the contradiction between productivity and data quantity requirements.
3Ease of manufacture
If manually tuned conventional controllers are used for deployment, then the controller can be deployed with existing techniques, but the process adds to costs and increases the time involved for robot deployment
Solution Approach 1:
The patent applies self-service by enabling the controller to automatically learn and adapt through reinforcement learning in simulation and real-world operation, eliminating the need for manual tuning by engineers. This self-service approach maintains ease of deployment while significantly reducing the time and cost associated with manual controller configuration.
Data Source
Figure 1
Figure 2
Figure 3
AI summary
Robotics control system (10) and method for training said robotics control system are provided. Disclosed embodiments make a gracefully blended utilization of Reinforcement Learning (RL) with conventional control by way of a dynamically adaptive interaction between respective control signals (20, 24) generated by a conventional feedback controller (18) and an RL controller (22). Additionally, disclosed embodiments make use of an iterative approach for training a control policy by effective use of virtual sensor and actuator data (60) interleaved with real-world sensor and actuator data (54). This is effective to reducing a training sample size to fulfill a blended control policy for the conventional feedback controller and the reinforcement learning controller. Disclosed embodiments may be used in a variety of industrial automation applications.