Robot Skill Learning With AR Policy Training and Real-World Feedback

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing robot learning methods face challenges in achieving accurate task skills due to long learning times, potential damage to task objects or robots during motion learning, and discrepancies between simulated and actual environments, leading to degraded accuracy when performing tasks in real environments.

Innovation Solution

A robot device equipped with a camera, robot arm, and control device that collects control records, implements an augmented reality model by rendering virtual objects in camera images, and performs image-based policy learning to update control policies, allowing for optimized task performance while minimizing errors and reducing risk in the actual work environment.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Reliability

If motion learning is performed using trial and error method in actual environment, then the robot can acquire optimized control policy, but learning time becomes long and risk of damage to task objects or robot increases

Engineering Contradiction:
Improvecontrol policy accuracyVSAvoidlearning time
Core Design Contradiction:
ReliabilityVSLoss of time

Solution Approach 1:

The patent performs preliminary motion learning in a simulated environment before deploying the robot in the actual environment. The simulation pre-trains the control policy using trial and error methods, so that when the robot operates in reality, it already possesses a baseline policy that requires minimal further adjustment, thereby reducing both learning time and risk of damage.

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The patent creates a virtual copy of the actual environment through simulation. This simulated environment replicates the physical space, task objects, and conditions, allowing the robot to learn control policies in a risk-free virtual setting before applying them in the real world, thus avoiding damage to actual objects during the learning process.

Inventive Principle:
Principle #26Copying

2Loss of time

If motion learning is replaced with simulation, then learning time is reduced and risk of damage is minimized, but accuracy is degraded due to differences between simulation and actual environment

Engineering Contradiction:
Improvelearning timeVSAvoidtask execution accuracy
Core Design Contradiction:
Loss of timeVSMeasurement precision

Solution Approach 1:

The patent implements a feedback mechanism where the robot performs tasks in the actual environment using the control policy learned from simulation, collects real-world task results and environmental data, and uses this feedback to fine-tune and adjust the control policy. This closed-loop approach corrects the discrepancies between simulation and reality, improving task execution accuracy while maintaining the time efficiency of simulation-based learning.

Inventive Principle:
Principle #23Feedback

3Productivity

If robot performs random motion for learning, then control policy can be optimized through reinforcement learning, but task objects or robot may be damaged when task fails

Engineering Contradiction:
Improvelearning efficiencyVSAvoiddamage risk to task objects
Core Design Contradiction:
ProductivityVSObject-affected harmful factors

Solution Approach 1:

The patent uses a simulated environment that copies the actual task space and objects. All random motions and trial-and-error learning occur in this virtual copy, allowing the robot to explore and learn control policies without any physical risk to real objects. The simulation absorbs all harmful effects of failed tasks while preserving learning efficiency.

Inventive Principle:
Principle #26Copying

4Productivity

If robot performs random motion for learning, then control policy can be optimized through reinforcement learning, but robot may collide with other parts within task space

Engineering Contradiction:
Improvelearning efficiencyVSAvoiddamage risk to robot
Core Design Contradiction:
ProductivityVSObject-generated harmful factors

Solution Approach 1:

The patent creates a virtual replica of the task space including all physical constraints and obstacles. The robot performs random motions and learning experiments in this simulated environment, where collisions with virtual objects have no physical consequences. This allows high-productivity learning while eliminating the risk of self-damage that would occur in the actual environment.

Inventive Principle:
Principle #26Copying

Data Source

PatentUS11911912B2Robot control apparatus and method for learning task skill of the robot
Publication Date: 2024.02.27 SAMSUNG ELECTRONICS CO LTD
  • US11911912B2 patent drawing
  • US11911912B2 patent drawing
  • US11911912B2 patent drawing

AI summary

A robot device according to various embodiments comprises a camera, a robot arm, and a control device electrically connected to the camera and the robot arm, wherein the control device can be configured to collect a robot arm control record about a random operation, acquire, from the camera, a camera image in which a working space of the robot arm is photographed, implement an augmented reality model by rendering a virtual object corresponding to an object related to objective work in the camera image, and update a control policy for the objective work by performing image-based policy learning on the basis of the augmented reality model and the control record.