Robot Control Policy Learning from Kinesthetic Teaching Data

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing techniques for robot control policies generated through kinesthetic teaching suffer from issues such as overweighting of data points in certain segments, leading to poorly conditioned learning and biased potential functions, and require tedious user input for potential gradients, which can impact performance and accuracy.

Innovation Solution

The method involves resampling temporally distributed data points to generate spatially distributed points, automatically determining potential gradients, and assigning prior weights and non-uniform smoothness parameters to improve the learning of non-parametric potential functions and dissipative fields, thereby generating a control policy that adapts to initial configurations and environmental changes.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Ease of manufacture

If temporally distributed data points are used directly for learning, then the data collection process is simple, but the learning becomes poorly conditioned and produces biased potential functions

Engineering Contradiction:
Improvedata collection simplicityVSAvoidlearning accuracy
Core Design Contradiction:
Ease of manufactureVSManufacturing precision

Solution Approach 1:

The patent transforms the parameter distribution of data points from temporal uniformity to spatial uniformity through resampling. This parameter transformation resolves the contradiction by changing how data points are distributed in space, ensuring that regions with higher robot velocities contribute appropriately to the potential function learning, thereby improving learning accuracy while maintaining the simplicity of data collection through kinesthetic teaching

Inventive Principle:
Principle #35Parameter changes

Solution Approach 2:

The patent applies preliminary processing to the data points by computing spatial distances and applying weighting factors before the actual learning process. This preliminary action of resampling and weighting ensures that the data is properly conditioned for learning, preventing the poorly conditioned learning problem while keeping the original data collection method intact

Inventive Principle:
Principle #10Preliminary action

2Ease of operation

If uniform temporal sampling is used during kinesthetic teaching, then the data collection is straightforward, but certain segments are overweighted leading to biased potential functions

Engineering Contradiction:
Improvedata collection easeVSAvoidpotential function accuracy
Core Design Contradiction:
Ease of operationVSReliability

Solution Approach 1:

The patent changes the sampling parameter from uniform temporal intervals to spatially uniform distribution. By computing the spatial distance between consecutive data points and using these distances to determine weighting factors, the patent ensures that each spatial region contributes equally to the potential function, eliminating the bias caused by velocity variations during kinesthetic teaching

Inventive Principle:
Principle #35Parameter changes

Solution Approach 2:

The patent incorporates feedback from the spatial distance calculations to adjust the weighting of data points. The computed spatial distances provide feedback about the robot's velocity and trajectory characteristics, which are then used to assign appropriate weights that correct for overweighting of certain segments, improving the reliability of the learned potential function

Inventive Principle:
Principle #23Feedback

3Ease of operation

If manual input of potential gradients is required, then user control over the process is maintained, but the process becomes tedious and performance is impacted

Engineering Contradiction:
Improveuser controlVSAvoidcontrol policy generation speed
Core Design Contradiction:
Ease of operationVSProductivity

Solution Approach 1:

The patent enables the system to automatically compute potential gradients from the demonstrated trajectories without requiring manual user input. The algorithm self-determines the gradients by analyzing the spatial distribution and weighting of data points, eliminating the tedious manual input process while maintaining reasonable user control through the kinesthetic teaching demonstration

Inventive Principle:
Principle #25Self-service

Solution Approach 2:

The patent replaces the manual mechanical process of specifying potential gradients with an automated computational process. Instead of requiring users to manually input gradient values, the system uses algorithmic computation based on the demonstrated trajectories, significantly improving productivity while still allowing users to influence the outcome through their physical demonstrations

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

4Reliability

If virtual data points are generated to ensure solution existence, then mathematical completeness is improved, but the complexity of the learning process increases

Engineering Contradiction:
Improvesolution existence guaranteeVSAvoidlearning process complexity
Core Design Contradiction:
ReliabilityVSDevice complexity

Solution Approach 1:

The patent extracts and addresses the core mathematical requirements directly from the actual demonstration data without adding virtual data points. By using proper weighting of real data points based on spatial distances, the method ensures solution existence while avoiding the complexity introduced by generating and processing virtual data points

Inventive Principle:
Principle #2Taking out (Extraction)

Data Source

PatentUS11554485B2Generating a robot control policy from demonstrations collected via kinesthetic teaching of a robot
Publication Date: 2023.01.17 GDM HOLDING LLC
  • US11554485B2 patent drawing
  • US11554485B2 patent drawing
  • US11554485B2 patent drawing

AI summary

Generating a robot control policy that regulates both motion control and interaction with an environment and/or includes a learned potential function and/or dissipative field. Some implementations relate to resampling temporally distributed data points to generate spatially distributed data points, and generating the control policy using the spatially distributed data points. Some implementations additionally or alternatively relate to automatically determining a potential gradient for data points, and generating the control policy using the automatically determined potential gradient. Some implementations additionally or alternatively relate to determining and assigning a prior weight to each of the data points of multiple groups, and generating the control policy using the weights. Some implementations additionally or alternatively relate to defining and using non-uniform smoothness parameters at each data point, defining and using d parameters for stiffness and/or damping at each data point, and/or obviating the need to utilize virtual data points in generating the control policy.