Robot Control Policy Training With AR Sensor Injection

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Training robot control policies using physical robots in real-world environments is costly and risky, while fully simulated environments fail to replicate real-world conditions, leading to inadequately trained policies.

Innovation Solution

Implementing augmented reality (AR) sensor data by injecting virtual objects into physical sensor data to train robot control policies, allowing physical robots to interact with virtual objects in controlled environments, thereby generating diverse training episodes with reduced costs and risks while simulating dynamic real-world scenarios.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Reliability

If training episodes are conducted using physical robots in real-world environments, then the robot control policy learns realistic physical phenomena and interactions, but the training process becomes prohibitively expensive and dangerous

Engineering Contradiction:
Improvetraining effectivenessVSAvoidrisk of harm to people and property
Core Design Contradiction:
ReliabilityVSObject-affected harmful factors

Solution Approach 1:

The patent creates virtual copies of real-world objects and environments that are injected into actual sensor data streams. These virtual objects replicate the visual, depth, and range characteristics of physical objects, allowing the robot to learn interactions with realistic phenomena without physical risk. The virtual objects are rendered to match camera perspectives, LIDAR depth maps, and other sensor modalities, creating a safe training environment that preserves learning value.

Inventive Principle:
Principle #26Copying

Solution Approach 2:

The system introduces virtual objects as an intermediary layer between the physical robot and the real-world environment. These virtual objects act as mediators that carry the essential interaction dynamics and physical phenomena the robot needs to learn, while eliminating the harmful aspects of real physical interactions. The virtual objects are integrated into sensor data pipelines, serving as a safe intermediate training medium.

Inventive Principle:
Principle #24Intermediary (Mediator)

2Object-affected harmful factors

If training episodes are conducted in controlled and lower entropy physical environments, then the risks are mitigated, but the environments do not realistically represent the real-world, resulting in inadequately trained robot control policy

Engineering Contradiction:
Improverisk mitigationVSAvoidtraining effectiveness
Core Design Contradiction:
Object-affected harmful factorsVSReliability

Solution Approach 1:

The system dynamically populates controlled physical environments with virtual objects that create diverse, high-entropy training scenarios. Instead of changing the physical environment, the system dynamically injects virtual pedestrians, objects, and obstacles into sensor data streams, allowing the robot to experience dynamic, unpredictable situations while remaining in a safe controlled setting. The virtual objects can be configured to simulate busy airports, festivals, or other complex environments.

Inventive Principle:
Principle #15Dynamics

Solution Approach 2:

The system changes the parameters of the training environment by injecting virtual objects with varying characteristics into sensor data. These virtual objects have different positions, velocities, sizes, and interaction properties that can be controlled and adjusted. This allows the training environment to transition from static and controlled to dynamic and realistic without changing the physical setting, effectively simulating high-entropy conditions in a safe manner.

Inventive Principle:
Principle #35Parameter changes

3Productivity

If wholly simulated robot episodes are implemented, then resources are significantly reduced and training can occur at much larger scale, but the simulated episodes fail to realistically represent the real world

Engineering Contradiction:
Improvetraining efficiencyVSAvoidrealism of training data
Core Design Contradiction:
ProductivityVSReliability

Solution Approach 1:

The patent merges the advantages of both physical and simulated training by combining real sensor data from physical robots with virtual objects from simulation. The system integrates virtual object rendering with actual camera feeds, LIDAR data, and other sensor streams, creating a hybrid training environment. This combination preserves the realism of physical sensor data while incorporating the scalability and diversity of virtual objects, achieving both efficiency and authenticity.

Inventive Principle:
Principle #5Merging (Combining)

Solution Approach 2:

The system creates virtual representations of objects and environments that are copied into the robot's actual sensor data streams. These virtual copies maintain the visual and spatial characteristics needed for realistic interaction learning, allowing the robot to process and respond to them as if they were real. The virtual objects are rendered to match the robot's sensor perspectives, preserving realism while enabling scalable training.

Inventive Principle:
Principle #26Copying

Data Source

PatentUS20240058954A1Training robot control policies
Publication Date: 2024.02.22 GDM HOLDING LLC
  • US20240058954A1 patent drawing
  • US20240058954A1 patent drawing
  • US20240058954A1 patent drawing

AI summary

Implementations are provided for training robot control policies using augmented reality (AR) sensor data comprising physical sensor data injected with virtual objects. In various implementations, physical pose(s) of physical sensor(s) of a physical robot operating in a physical environment may be determined. Virtual pose(s) of virtual object(s) in the physical environment may also be determined. Based on the physical poses virtual poses, the virtual object(s) may be injected into sensor data generated by the one or more physical sensors to generate AR sensor data. The physical robot may be operated in the physical environment based on the AR sensor data and a robot control policy. The robot control policy may be trained based on virtual interactions between the physical robot and the one or more virtual objects.