Robot Control Policy Learning With Contracting Vector Fields

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing robot control techniques, such as time-indexed trajectory methods and dynamical systems approaches, face inefficiencies in time required to traverse trajectories, robot wear, and inability to handle dynamic obstacles effectively, with non-convex optimization leading to sub-optimal results.

Innovation Solution

The use of vector-valued Reproducing Kernel Hilbert spaces (RKHS), contraction analysis, and convex optimization to learn stable, non-linear dynamical systems for robot control, generating vector fields that induce a contraction tube around learned trajectories, allowing for adaptation to dynamic environments and efficient policy generation.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Ease of manufacture

If time-indexed trajectory methods are used to control robot motion, then the trajectory can be easily recorded and reproduced, but the time required to traverse the trajectory increases and robot wear increases

Engineering Contradiction:
Improveease of trajectory recordingVSAvoidtrajectory traversal time
Core Design Contradiction:
Ease of manufactureVSLoss of time

Solution Approach 1:

The patent transforms the static time-indexed trajectory into a dynamic system by formulating robot motion as differential equations with state variables that evolve over time. This allows the system to adapt its motion profile dynamically, optimizing traversal speed while maintaining accuracy and reducing wear through intelligent velocity and acceleration control.

Inventive Principle:
Principle #15Dynamics

Solution Approach 2:

The patent changes the fundamental parameters of trajectory representation from time-indexed waypoints to dynamical system parameters (vector fields, equilibrium points, and invariant manifolds). This parameter transformation enables the robot to compute optimal trajectories analytically rather than following pre-recorded paths, significantly reducing traversal time and mechanical stress.

Inventive Principle:
Principle #35Parameter changes

2Ease of operation

If simple control policies are used for trajectory following, then the control implementation is straightforward, but the robot cannot effectively deal with dynamic obstacles in the environment

Engineering Contradiction:
Improvecontrol policy simplicityVSAvoidadaptability to dynamic obstacles
Core Design Contradiction:
Ease of operationVSAdaptability or versatility

Solution Approach 1:

The patent employs dynamic systems theory to create control policies that are inherently adaptive. By representing desired trajectories as invariant manifolds of a dynamical system, the controller can naturally respond to disturbances and dynamic obstacles while maintaining stability, combining mathematical rigor with environmental adaptability.

Inventive Principle:
Principle #15Dynamics

Solution Approach 2:

The patent incorporates feedback mechanisms through the dynamical system formulation, where the robot continuously monitors its state and adjusts its motion to remain on the desired invariant manifold. This feedback enables real-time adaptation to dynamic obstacles while maintaining the simplicity of the underlying control structure.

Inventive Principle:
Principle #23Feedback

3Adaptability or versatility

If dynamical systems approaches with non-convex optimization are used to learn control policies, then the policy can capture essential dynamics and adapt to changes, but the optimization is prone to sub-optimal local minima

Engineering Contradiction:
Improvegeneralization capabilityVSAvoidoptimization convergence reliability
Core Design Contradiction:
Adaptability or versatilityVSReliability

Solution Approach 1:

The patent extracts and separates the optimization problem into distinct components: the dynamical system structure is predefined based on task requirements, and only specific parameters (equilibrium points, vector field coefficients) need to be learned from demonstrations. This decomposition transforms the complex non-convex optimization into a more manageable problem with guaranteed convergence properties.

Inventive Principle:
Principle #2Taking out (Extraction)

Solution Approach 2:

The patent performs preliminary action by pre-defining the dynamical system structure and constraints before optimization. By establishing the mathematical framework and invariant manifold structure in advance, the subsequent learning process only needs to fit specific parameters, avoiding the pitfalls of fully non-convex optimization while maintaining generalization capability.

Inventive Principle:
Principle #10Preliminary action

Data Source

PatentUS11565412B2Generating a robot control policy from demonstrations collected via kinesthetic teaching of a robot
Publication Date: 2023.01.31 GOOGLE LLC
  • US11565412B2 patent drawing
  • US11565412B2 patent drawing
  • US11565412B2 patent drawing

AI summary

Techniques are described herein for generating a dynamical systems control policy. A non-parametric family of smooth maps is defined on which vector-field learning problems can be formulated and solved using convex optimization. In some implementations, techniques described herein address the problem of generating contracting vector fields for certifying stability of the dynamical systems arising in robotics applications, e.g., designing stable movement primitives. These learning problems may utilize a set of demonstration trajectories, one or more desired equilibria (e.g., a target point), and once or more statistics including at least an average velocity and average duration of the set of demonstration trajectories. The learned contracting vector fields may induce a contraction tube around a targeted trajectory for an end effector of the robot. In some implementations, the disclosed framework may use curl-free vector-valued Reproducing Kernel Hilbert Spaces.