Kinesthetic Robot Policy Learning with Spatial Resampling and Weights

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing methods for generating robot control policies from kinesthetic teaching data face challenges such as overweighting of data points in certain segments, leading to biased learning and poor performance, especially when dealing with changing environmental conditions and adapting motion in real-time.

Innovation Solution

The approach involves resampling temporally distributed data points to generate spatially distributed points, automatically determining potential gradients, assigning prior weights, and defining non-uniform smoothness parameters to improve the learning of a non-parametric potential function and dissipative field, thereby generating a control policy that adapts to initial configurations and environmental changes.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Ease of manufacture

If temporally distributed data points are used directly for learning potential functions, then the learning process is simpler, but the control policy performance deteriorates due to overweighting of data points in certain segments

Engineering Contradiction:
Improvesimplicity of learning processVSAvoidcontrol policy performance
Core Design Contradiction:
Ease of manufactureVSReliability

Solution Approach 1:

The patent transforms the distribution parameter of data points from temporal to spatial domain. By resampling data points based on spatial intervals rather than temporal intervals, the method changes the weighting parameter distribution to achieve uniform representation across all trajectory segments, eliminating the overweighting issue while maintaining learning effectiveness

Inventive Principle:
Principle #35Parameter changes

Solution Approach 2:

The patent creates a copy of the temporal data sequence and transforms it into spatially distributed data points. This copying and transformation process allows the system to preserve the original temporal information while generating a new spatial representation that provides balanced weighting for potential function learning

Inventive Principle:
Principle #26Copying

2Adaptability or versatility

If the control policy is designed to adapt to changing environmental conditions and initial configurations, then the versatility improves, but the complexity of the learning process increases

Engineering Contradiction:
Improveadaptability to environmental changesVSAvoidcomplexity of learning process
Core Design Contradiction:
Adaptability or versatilityVSDevice complexity

Solution Approach 1:

The patent develops a universal learning framework that handles multiple scenarios (different environmental conditions, various initial configurations, different task types) through a single spatially distributed data point approach. This universal method eliminates the need for separate learning processes for different conditions, thereby reducing overall system complexity while maintaining high adaptability

Inventive Principle:
Principle #6Universality (Multi-functionality)

Solution Approach 2:

The patent uses parameter transformation (from temporal to spatial distribution) as a universal solution that works across different environmental conditions and task types. This parameter change approach provides a consistent learning mechanism that adapts to various scenarios without requiring complex condition-specific adjustments

Inventive Principle:
Principle #35Parameter changes

Data Source

PatentEP3795307B1Improvements related to generating a robot control policy from demonstrations collected via kinesthetic teaching of a robot
Publication Date: 2022.09.28 X DEVELOPMENT LLC
  • EP3795307B1 patent drawingFigure 1~2
  • EP3795307B1 patent drawingFigure 3
  • EP3795307B1 patent drawingFigure 4

AI summary

The application relates to a method for generating a robot control policy that regulates both motion control and interaction with an environment and which includes a learned potential function and optionally a dissipative field. The method comprises the steps of automatically determining a potential gradient for data points and generating the control policy using the automatically determined potential gradient. Some implementations relate to resampling temporally distributed data points to generate spatially distributed data points, and generating the control policy using the spatially distributed data points. Some implementations additionally or alternatively relate to determining and assigning a prior weight to each of the data points of multiple groups, and generating the control policy using the weights. Some implementations additionally or alternatively relate to defining and using non-uniform smoothness parameters at each data point, defining and using d parameters for stiffness and/or damping at each data point, and/or obviating the need to utilize virtual data points in generating the control policy.