Robot Control Policy Learning for Dynamic Obstacle Adaptation

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing robot control techniques, such as time-indexed trajectory methods, fail to optimize time required for trajectory traversal and robot wear, and are inadequate for handling dynamic obstacles in unstructured environments.

Innovation Solution

A dynamical systems approach using contraction theory and semidefinite programming to generate a polynomial contracting vector field (CVF-P) from demonstration data, enabling real-time adaptation to dynamic obstacles and ensuring continuous-time guarantees on imitation behavior.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Productivity

If time-indexed trajectory methods are used to control robot motion, then the trajectory can be recorded and repeated, but the time required to traverse the trajectory is not optimized and robot wear increases

Engineering Contradiction:
Improvetrajectory traversal timeVSAvoidrobot wear and tear
Core Design Contradiction:
ProductivityVSLoss of time

Solution Approach 1:

The patent transforms the static time-indexed trajectory into a dynamic system using differential equations. The robot motion is governed by a learned dynamical system that adapts to different starting positions and optimizes the traversal path in real-time, rather than rigidly following pre-recorded waypoints. This dynamic approach allows the robot to minimize traversal time and wear by continuously adjusting its motion based on the underlying task dynamics.

Inventive Principle:
Principle #15Dynamics

2Adaptability or versatility

If simple control policies are used for trajectory following, then the control implementation is straightforward, but the robot cannot adapt to dynamic obstacles in the environment

Engineering Contradiction:
Improveadaptation to dynamic obstaclesVSAvoidcontrol policy complexity
Core Design Contradiction:
Adaptability or versatilityVSDevice complexity

Solution Approach 1:

The patent implements feedback control through the learned dynamical system. The system continuously monitors the robot's state and environmental conditions, then adjusts the control inputs accordingly. The differential equation-based model learns from demonstrations and incorporates feedback loops that enable real-time adaptation to dynamic obstacles while maintaining stable and predictable robot behavior. This feedback mechanism allows the robot to deviate from the original trajectory when necessary and return to the task goal.

Inventive Principle:
Principle #23Feedback

3Adaptability or versatility

If demonstration trajectories are memorized by the policy, then the robot can reproduce the demonstrated behavior, but the policy lacks generalization to new situations

Engineering Contradiction:
Improvegeneralization capabilityVSAvoidimitation accuracy
Core Design Contradiction:
Adaptability or versatilityVSReliability

Solution Approach 1:

The patent transforms the control approach by changing from memorizing fixed trajectory parameters to learning continuous dynamical system parameters. Instead of storing discrete waypoints and timing information, the system learns the underlying differential equations that govern the task dynamics. This allows the policy to generate appropriate trajectories for novel starting positions and environmental conditions while maintaining faithful reproduction of the demonstrated task behavior through the learned dynamics.

Inventive Principle:
Principle #35Parameter changes

Data Source

PatentEP3898132B1Generating a robot control policy from demonstrations
Publication Date: 2025.08.20 GOOGLE LLC
  • EP3898132B1 patent drawingFigure 1~2
  • EP3898132B1 patent drawingFigure 3~4
  • EP3898132B1 patent drawingFigure 5A~5B

AI summary

Learning to effectively imitate human teleoperators, even in unseen, dynamic environments is a promising path to greater autonomy, enabling robots to steadily acquire complex skills from supervision. Various motion generation techniques are described herein that are rooted in contraction theory and sum-of-squares programming for learning a dynamical systems control policy in the form of a polynomial vector field from a given set of demonstrations. Notably, this vector field is provably optimal for the problem of minimizing imitation loss while providing certain continuous-time guarantees on the induced imitation behavior. Techniques herein generalize to new initial and goal poses of the robot and can adapt in real time to dynamic obstacles during execution, with convergence to teleoperator behavior within a well-defined safety tube.