Control Agent Training With Decaying Supervisor Guidance

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Conventional control methods for robots and systems are inefficient and unsafe, particularly in complex environments, due to the need for manually designed models and large amounts of training data, which can lead to costly and hazardous errors during training.

Innovation Solution

The training of a learning agent is alternated with a pioneer agent under the supervision of a supervisor agent, with a decaying supervisor coefficient to reduce the influence of the supervisor, allowing for efficient and safe real-time control by iteratively updating policies and reducing the risk of costly errors.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Adaptability or versatility

If conventional reinforcement learning techniques are used to explore unknown environmental state spaces, then the control system can learn to perform tasks, but the computation time becomes unwieldy and catastrophic failures or hazardous errors occur frequently during early learning stages

Engineering Contradiction:
Improveability to explore unknown environmental state spacesVSAvoidcomputation time for adequate exploration
Core Design Contradiction:
Adaptability or versatilityVSLoss of time

Solution Approach 1:

The patent applies preliminary action by pre-training a supervisor agent in a simulated environment before deploying it to supervise the learning agent in the physical environment. This preliminary training in simulation allows the supervisor to develop competent policies without risking costly errors in the real world, thereby reducing computation time and hazards during actual exploration.

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The supervisor agent acts as an intermediary between the learning agent and the physical environment. It supervises and guides the learning agent's actions, filtering out hazardous explorations while allowing beneficial learning to occur. This intermediary role reduces the computation time needed for safe exploration by preventing catastrophic failures during the learning process.

Inventive Principle:
Principle #24Intermediary (Mediator)

2Reliability

If the speed and magnitude of actions of the controlled object are decreased during training to avoid costly errors, then safety is improved, but the training fails to converge to safe and effective control agents within acceptable training times

Engineering Contradiction:
Improvesafety during trainingVSAvoidtraining convergence speed
Core Design Contradiction:
ReliabilityVSProductivity

Solution Approach 1:

The supervisor agent is pre-trained in simulation before being deployed to supervise real-world training. This preliminary action allows it to learn safe and effective control policies in advance, enabling it to guide the learning agent at full speed and magnitude without causing costly errors, thus maintaining both safety and training productivity.

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The supervisor agent learns by copying successful strategies from simulation to the physical environment. Instead of slowly exploring from scratch in the real world, it replicates competent behaviors developed in the virtual environment, achieving rapid convergence to safe and effective control policies without the need for slow, cautious real-world experimentation.

Inventive Principle:
Principle #26Copying

3Reliability

If a simulated environment is developed to train the control agent and avoid costly errors, then training safety is improved, but the computational time required becomes unacceptably large and differences between simulated and physical environments may be too great

Engineering Contradiction:
Improvetraining safety in simulated environmentVSAvoidcomputational time for simulation training
Core Design Contradiction:
ReliabilityVSLoss of time

Solution Approach 1:

The supervisor agent is trained in a simulated environment and then copies its learned policies to supervise the learning agent in the physical environment. This copying approach allows most of the training to occur quickly in simulation, with only minimal real-world supervision needed, thereby reducing overall computational time while maintaining training safety.

Inventive Principle:
Principle #26Copying

Solution Approach 2:

The simulated environment training is performed as a preliminary step before physical deployment. The supervisor agent develops competent policies in simulation, and this preliminary training reduces the need for extensive real-world experimentation, thereby reducing the total computational time required while ensuring training safety.

Inventive Principle:
Principle #10Preliminary action

Data Source

PatentUS11709462B2Safe and efficient training of a control agent
Publication Date: 2023.07.25 ADOBE INC
  • US11709462B2 patent drawing
  • US11709462B2 patent drawing
  • US11709462B2 patent drawing

AI summary

The training of a learning agent to provide real-time control of an object is disclosed. Training of the learning agent and training of a corresponding pioneer agent are iteratively alternated. The training of the learning and pioneer agents is under the supervision of a supervisor agent. The training of the learning agent provides feedback for subsequent training of the pioneer agent. The training of the pioneer agent provides feedback for subsequent training of the learning agent. During the training, a supervisor coefficient modulates the influence of the supervisor agent. As agents are trained, the influence of the supervisor agent is decayed. The training of the learning agent, under a first level of supervisor influence, includes real-time control of the object. The subsequent training of the pioneer agent, under a reduced level of supervisor influence, includes replay of training data accumulated during the real-time control of the object.