Real-Time Reinforcement Learning Architecture for Delayed Robot Control

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Reinforcement learning models face inefficiencies in real-world environments due to issues like slow data collection, partial observability, noisy sensors, and time delays, which hinder their performance in controlling physical robotic devices.

Innovation Solution

A reinforcement learning architecture that includes device communicators, a task manager, and a reinforcement learning agent operating independently, collecting and processing joint state vectors to reduce time delays and improve decision consistency, allowing for adaptive control of multiple devices in real-time environments.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Adaptability or versatility

If reinforcement learning models are applied to real-world environments with physical robotic devices, then the system can achieve practical control tasks, but the performance deteriorates due to time delays and data collection issues

Engineering Contradiction:
Improvereal-world applicabilityVSAvoidlearning performance
Core Design Contradiction:
Adaptability or versatilityVSReliability

Solution Approach 1:

The patent creates a virtual copy of the physical environment through high-fidelity simulation. The simulation replicates the dynamics, sensors, and actuators of the physical robotic system, allowing the reinforcement learning agent to train extensively in the virtual environment before deployment to the real world, thereby overcoming time delay and data collection limitations

Inventive Principle:
Principle #26Copying

Solution Approach 2:

The patent introduces a simulation environment as an intermediary between the reinforcement learning agent and the physical robotic system. This intermediary layer allows for rapid experimentation and learning without the constraints of physical time delays, while still maintaining fidelity to the real-world system dynamics

Inventive Principle:
Principle #24Intermediary (Mediator)

2Speed

If reinforcement learning operates in real-time with physical devices, then practical control is achieved, but time delays between sensorimotor packets and action commands deteriorate learning performance

Engineering Contradiction:
Improvereal-time control speedVSAvoidtime delay between action and observation
Core Design Contradiction:
SpeedVSLoss of time

Solution Approach 1:

The patent performs preliminary training and action generation in the simulation environment where time delays are negligible. The reinforcement learning agent learns policies and generates actions in advance within the virtual environment, allowing extensive exploration and learning without the constraint of physical time delays before deploying to the real system

Inventive Principle:
Principle #10Preliminary action

3Loss of information

If reinforcement learning systems collect data from physical sensors in real-time, then real-world learning is achieved, but slow data collection rate and noisy sensors reduce learning efficiency

Engineering Contradiction:
Improvesensor data qualityVSAvoiddata collection rate
Core Design Contradiction:
Loss of informationVSProductivity

Solution Approach 1:

The patent replicates the sensor model and noise characteristics in the simulation environment. The virtual sensors produce data with the same statistical properties and noise patterns as the physical sensors, allowing the agent to learn from large volumes of clean, rapidly generated simulation data that accurately reflects real-world sensor behavior

Inventive Principle:
Principle #26Copying

Data Source

PatentUS12005578B2Real-time real-world reinforcement learning systems and methods
Publication Date: 2024.06.11 OCADO INNOVATION LTD
  • US12005578B2 patent drawing
  • US12005578B2 patent drawing
  • US12005578B2 patent drawing

AI summary

A reinforcement learning architecture for facilitating reinforcement learning in connection with operation of an external real-time system that includes a plurality of devices operating in a real-world environment. The reinforcement learning architecture includes a plurality of communicators, a task manager, and a reinforcement learning agent that interact with each other to effectuate a policy for achieving a defined objective in the real-world environment. Each of the communicators receives sensory data from a corresponding device and the task manager generates a joint state vector based on the sensory data. The reinforcement learning agent generates, based on the joint state vector, a joint action vector, which the task manager parses into a plurality of actuation commands. The communicators transmit the actuation commands to the plurality of devices in the real-world environment.