Robot Control Policy Learning from Spatially Resampled Demonstrations

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing techniques for teaching robots through kinesthetic methods face challenges in generating effective control policies that adapt to changing environmental conditions and ensure smooth motion, often resulting in biased learning due to uneven data point distribution and the need for virtual data points.

Innovation Solution

The method involves resampling temporally distributed data points to generate spatially distributed points, learning a non-parametric potential function, and determining potential gradients and prior weights to create a control policy that regulates robot motion and interaction, while eliminating the need for virtual data points.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Reliability

If temporally distributed data points are used directly for learning control policy, then the learning process is simpler, but the control policy performance is degraded due to biased learning from uneven data distribution

Engineering Contradiction:
Improvecontrol policy performanceVSAvoiddata processing complexity
Core Design Contradiction:
ReliabilityVSDevice complexity

Solution Approach 1:

The patent applies preliminary action by resampling temporally distributed data points into spatially distributed data points before the learning process. This preprocessing step ensures uniform spatial distribution of data points, eliminating bias in the learning process and improving control policy performance without adding significant complexity to the overall system.

Inventive Principle:
Principle #10Preliminary action

2Measurement precision

If virtual data points are generated to improve data distribution, then learning accuracy is improved, but computational complexity and processing time increase

Engineering Contradiction:
Improvelearning accuracyVSAvoidcomputational time
Core Design Contradiction:
Measurement precisionVSLoss of time

Solution Approach 1:

The patent uses copying by resampling existing real data points to create spatially distributed representations rather than generating virtual data points. This approach maintains learning accuracy by ensuring uniform spatial coverage while avoiding the computational overhead of generating and processing synthetic virtual data points.

Inventive Principle:
Principle #26Copying

3Productivity

If non-uniform data point distribution is accepted, then data collection is faster, but the control policy fails to adapt to changing environmental conditions

Engineering Contradiction:
Improvedata collection speedVSAvoidenvironmental adaptation
Core Design Contradiction:
ProductivityVSAdaptability or versatility

Solution Approach 1:

The patent applies preliminary action by transforming the collected data from temporal to spatial distribution before learning. This transformation enables the control policy to adapt to changing environmental conditions by ensuring uniform spatial coverage of data points, while the actual data collection process remains fast and uninterrupted.

Inventive Principle:
Principle #10Preliminary action

4Manufacturing precision

If more data points are collected to improve learning quality, then control accuracy is improved, but the computational burden and processing complexity increase

Engineering Contradiction:
Improvecontrol accuracyVSAvoidprocessing complexity
Core Design Contradiction:
Manufacturing precisionVSDevice complexity

Solution Approach 1:

The patent applies preliminary action by resampling data points to achieve uniform spatial distribution before the learning process. This preprocessing step improves control accuracy by eliminating sampling bias, while actually reducing processing complexity by creating a more efficient and balanced dataset for learning algorithms.

Inventive Principle:
Principle #10Preliminary action

Data Source

PatentUS11872699B2Generating a robot control policy from demonstrations collected via kinesthetic teaching of a robot
Publication Date: 2024.01.16 GDM HOLDING LLC
  • US11872699B2 patent drawing
  • US11872699B2 patent drawing
  • US11872699B2 patent drawing

AI summary

Generating a robot control policy that regulates both motion control and interaction with an environment and/or includes a learned potential function and/or dissipative field. Some implementations relate to resampling temporally distributed data points to generate spatially distributed data points, and generating the control policy using the spatially distributed data points. Some implementations additionally or alternatively relate to automatically determining a potential gradient for data points, and generating the control policy using the automatically determined potential gradient. Some implementations additionally or alternatively relate to determining and assigning a prior weight to each of the data points of multiple groups, and generating the control policy using the weights. Some implementations additionally or alternatively relate to defining and using non-uniform smoothness parameters at each data point, defining and using d parameters for stiffness and/or damping at each data point, and/or obviating the need to utilize virtual data points in generating the control policy.