Adaptive Controller Combining Reinforcement and Supervised Learning
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing robotic devices face challenges in adapting to changes in their model or environment, as programming is costly and remote control requires human operators, which can be inadequate in dynamic situations, and may not efficiently handle unexpected obstacles.
Innovation Solution
An adaptive controller system that combines reinforcement learning and supervised learning processes using a predictor and combiner, generating a control output based on sensory input and teaching signals, allowing the robotic apparatus to execute maneuvers autonomously.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Adaptability or versatility
If robotic devices are programmed to perform desired functionality, then the robot can execute specific tasks, but changes in the robot model or environment require changes in the programming code which is costly
Solution Approach 1:
The robotic device performs self-learning through reinforcement learning and supervised learning processes. The predictor component learns to generate predicted control outputs autonomously by receiving teaching signals from the combiner, eliminating the need for external reprogramming when environmental changes occur. The system serves itself by automatically adapting its control strategy based on performance feedback.
Solution Approach 2:
The predictor is pre-configured with the capability to learn through supervised learning using teaching signals. Instead of reprogramming the entire system when changes are needed, the predictor preliminarily acquires learning capacity during initial setup, enabling it to rapidly adapt to new situations by learning from combined control outputs without costly reprogramming interventions.
2Adaptability or versatility
If robotic devices are remotely controlled by humans, then the robot can handle complex tasks requiring user experience, but human operators are required and may be inadequate when dynamics change rapidly
Solution Approach 1:
The system implements a feedback loop where the combiner receives the control output from the adaptive controller and generates teaching signals based on the difference between actual and desired performance. This feedback mechanism enables the predictor to continuously learn and improve its predictions, allowing the robotic device to autonomously adapt to rapid dynamic changes without human intervention.
Solution Approach 2:
The patent replaces the mechanical system of human remote control with an automated learning system. The predictor, through supervised learning using teaching signals from the combiner, substitutes human operators by autonomously generating accurate control outputs in response to rapid environmental changes, eliminating the limitations of human reaction time and availability.
3Adaptability or versatility
If robotic devices learn to operate via exploration, then the robot can adapt to new situations, but the learning process is slow and inefficient
Solution Approach 1:
The system merges reinforcement learning and supervised learning processes by combining the control output from the adaptive controller with predicted control outputs from the predictor. The combiner integrates these two learning approaches, allowing the system to benefit from both the exploratory nature of reinforcement learning and the directed efficiency of supervised learning, significantly reducing the time required to learn new operations compared to exploration alone.
Solution Approach 2:
The combiner acts as an intermediary that bridges reinforcement learning and supervised learning. It receives the control output from reinforcement learning and generates teaching signals for supervised learning based on performance differences. This intermediary mechanism enables efficient knowledge transfer and acceleration of the learning process by guiding the predictor toward optimal solutions rather than relying solely on slow trial-and-error exploration.
Data Source
AI summary
Framework may be implemented for transferring knowledge from an external agent to a robotic controller. In an obstacle avoidance/target approach application, the controller may be configured to determine a teaching signal based on a sensory input, the teaching signal conveying information associated with target action consistent with the sensory input, the sensory input being indicative of the target/obstacle. The controller may be configured to determine a control signal based on the sensory input, the control signal conveying information associated with target approach/avoidance action. The controller may determine a predicted control signal based on the sensory input and the teaching signal, the predicted control conveying information associated with the target action. The control signal may be combined with the predicted control in order to cause the robotic apparatus to execute the target action.


