Robotic Manipulator Model Learning Without Velocity Sensors

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

The application of reinforcement learning to real physical systems, such as robotic systems, is challenging due to the need for extensive experience and safety risks associated with random exploration, and existing model learning methods require measurements of velocities and accelerations, which are often unavailable, leading to inaccurate predictions.

Innovation Solution

A derivative-free model learning framework that uses a finite past history of position measurements to represent the system state, eliminating the need for velocity and acceleration measurements, and employs semi-parametric Gaussian Process Regression models to improve generalization and prediction accuracy.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Extent of automation

If reinforcement learning is applied to real physical systems, then intelligent decision-making capability is improved, but safety risks increase due to random exploration requirements

Engineering Contradiction:
Improveintelligent decision-making capabilityVSAvoidsafety
Core Design Contradiction:
Extent of automationVSReliability

Solution Approach 1:

The patent applies preliminary action by pre-training a simulation model using reinforcement learning before deploying it to the real physical system. The simulation environment allows the agent to learn optimal control policies through random exploration without risking damage to the actual robotic system. Once trained in simulation, the learned policy is transferred to the real system, eliminating the need for dangerous random exploration in the physical environment.

Inventive Principle:
Principle #10Preliminary action

2Measurement precision

If traditional model learning methods are used that require velocity and acceleration measurements, then prediction accuracy may be improved, but device complexity increases due to additional sensor requirements

Engineering Contradiction:
Improveprediction accuracyVSAvoidsensor requirements
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The patent applies the taking out principle by extracting and eliminating the requirement for velocity and acceleration sensors from the system. The method reformulates the dynamic model to depend only on position measurements, which are already available from standard encoder sensors. This extraction of the problematic measurement requirements simplifies the device while maintaining prediction accuracy through a restructured modeling approach that uses position-only data with augmented state variables.

Inventive Principle:
Principle #2Taking out (Extraction)

3Adaptability or versatility

If Gaussian Process Regression models are used for model-based reinforcement learning, then generalization capability is improved, but computational load increases

Engineering Contradiction:
Improvegeneralization capabilityVSAvoidcomputational load
Core Design Contradiction:
Adaptability or versatilityVSPower

Solution Approach 1:

The patent applies segmentation by dividing the computational task into distinct phases: offline training phase where the Gaussian Process model is trained using simulation data, and online execution phase where the trained model is used for prediction. The complex GP training computations are performed offline when computational resources are abundant, while the online phase requires minimal computation for real-time control decisions. This temporal segmentation reduces the immediate computational load during robot operation while maintaining high generalization capability.

Inventive Principle:
Principle #1Segmentation

Data Source

PatentUS11389957B2System and design of derivative-free model learning for robotic systems
Publication Date: 2022.07.19 MITSUBISHI ELECTRIC RESEARCH LABORATORIES INC
  • US11389957B2 patent drawing
  • US11389957B2 patent drawing
  • US11389957B2 patent drawing

AI summary

A manipulator learning-control apparatus for controlling a manipulating system that includes an interface configured to receive manipulator state signals of the manipulating system and object state signals with respect to an object to be manipulated by the manipulating system in a workspace, wherein the object state signals are detected by at least one object detector, an output interface configured to transmit initial and updated policy programs to the manipulating system, a memory to store computer-executable programs including a data preprocess program, object state history data, manipulator state history data, a Derivative-Free Semi-parametric Gaussian Process (DF-SPGP) kernel learning program, a Derivative-Free Semi-parametric Gaussian Process (DF-SPGP) model learning program, an update-policy program and an initial policy program, and a processor, in connection with the memory, configured to transmit the initial policy program to the manipulating system for initiating a learning process that operates the manipulator system manipulating the object while a preset period of time.