Adaptive Controller Combining Reinforcement and Supervised Learning

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing robotic devices face challenges in adapting to changes in their model or environment, as programming is costly and remote control requires human operators, which can be inadequate in dynamic situations, and may not efficiently handle unexpected obstacles.

Innovation Solution

An adaptive controller system that combines reinforcement learning and supervised learning processes using a predictor and combiner, generating a control output based on sensory input and teaching signals, allowing the robotic apparatus to execute maneuvers autonomously.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Adaptability or versatility

If robotic devices are programmed to perform desired functionality, then the robot can execute specific tasks, but changes in the robot model or environment require changes in the programming code which is costly

Engineering Contradiction:
Improveadaptability to environmental changesVSAvoidcost of reprogramming
Core Design Contradiction:
Adaptability or versatilityVSEase of manufacture

Solution Approach 1:

The robotic device performs self-learning through reinforcement learning and supervised learning processes. The predictor component learns to generate predicted control outputs autonomously by receiving teaching signals from the combiner, eliminating the need for external reprogramming when environmental changes occur. The system serves itself by automatically adapting its control strategy based on performance feedback.

Inventive Principle:
Principle #25Self-service

Solution Approach 2:

The predictor is pre-configured with the capability to learn through supervised learning using teaching signals. Instead of reprogramming the entire system when changes are needed, the predictor preliminarily acquires learning capacity during initial setup, enabling it to rapidly adapt to new situations by learning from combined control outputs without costly reprogramming interventions.

Inventive Principle:
Principle #10Preliminary action

2Adaptability or versatility

If robotic devices are remotely controlled by humans, then the robot can handle complex tasks requiring user experience, but human operators are required and may be inadequate when dynamics change rapidly

Engineering Contradiction:
Improveresponse to dynamic changesVSAvoidautonomous operation capability
Core Design Contradiction:
Adaptability or versatilityVSExtent of automation

Solution Approach 1:

The system implements a feedback loop where the combiner receives the control output from the adaptive controller and generates teaching signals based on the difference between actual and desired performance. This feedback mechanism enables the predictor to continuously learn and improve its predictions, allowing the robotic device to autonomously adapt to rapid dynamic changes without human intervention.

Inventive Principle:
Principle #23Feedback

Solution Approach 2:

The patent replaces the mechanical system of human remote control with an automated learning system. The predictor, through supervised learning using teaching signals from the combiner, substitutes human operators by autonomously generating accurate control outputs in response to rapid environmental changes, eliminating the limitations of human reaction time and availability.

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

3Adaptability or versatility

If robotic devices learn to operate via exploration, then the robot can adapt to new situations, but the learning process is slow and inefficient

Engineering Contradiction:
Improvelearning capabilityVSAvoidlearning time
Core Design Contradiction:
Adaptability or versatilityVSLoss of time

Solution Approach 1:

The system merges reinforcement learning and supervised learning processes by combining the control output from the adaptive controller with predicted control outputs from the predictor. The combiner integrates these two learning approaches, allowing the system to benefit from both the exploratory nature of reinforcement learning and the directed efficiency of supervised learning, significantly reducing the time required to learn new operations compared to exploration alone.

Inventive Principle:
Principle #5Merging (Combining)

Solution Approach 2:

The combiner acts as an intermediary that bridges reinforcement learning and supervised learning. It receives the control output from reinforcement learning and generates teaching signals for supervised learning based on performance differences. This intermediary mechanism enables efficient knowledge transfer and acceleration of the learning process by guiding the predictor toward optimal solutions rather than relying solely on slow trial-and-error exploration.

Inventive Principle:
Principle #24Intermediary (Mediator)

Data Source

PatentUS9008840B1Apparatus and methods for reinforcement-guided supervised learning
Publication Date: 2015.04.14 BRAIN CORP
  • US9008840B1 patent drawing
  • US9008840B1 patent drawing
  • US9008840B1 patent drawing

AI summary

Framework may be implemented for transferring knowledge from an external agent to a robotic controller. In an obstacle avoidance/target approach application, the controller may be configured to determine a teaching signal based on a sensory input, the teaching signal conveying information associated with target action consistent with the sensory input, the sensory input being indicative of the target/obstacle. The controller may be configured to determine a control signal based on the sensory input, the control signal conveying information associated with target approach/avoidance action. The controller may determine a predicted control signal based on the sensory input and the teaching signal, the predicted control conveying information associated with the target action. The control signal may be combined with the predicted control in order to cause the robotic apparatus to execute the target action.