Robotic Manipulator Model Learning Without Velocity Sensors
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
The application of reinforcement learning to real physical systems, such as robotic systems, is challenging due to the need for extensive experience and safety risks associated with random exploration, and existing model learning methods require measurements of velocities and accelerations, which are often unavailable, leading to inaccurate predictions.
Innovation Solution
A derivative-free model learning framework that uses a finite past history of position measurements to represent the system state, eliminating the need for velocity and acceleration measurements, and employs semi-parametric Gaussian Process Regression models to improve generalization and prediction accuracy.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Extent of automation
If reinforcement learning is applied to real physical systems, then intelligent decision-making capability is improved, but safety risks increase due to random exploration requirements
Solution Approach 1:
The patent applies preliminary action by pre-training a simulation model using reinforcement learning before deploying it to the real physical system. The simulation environment allows the agent to learn optimal control policies through random exploration without risking damage to the actual robotic system. Once trained in simulation, the learned policy is transferred to the real system, eliminating the need for dangerous random exploration in the physical environment.
2Measurement precision
If traditional model learning methods are used that require velocity and acceleration measurements, then prediction accuracy may be improved, but device complexity increases due to additional sensor requirements
Solution Approach 1:
The patent applies the taking out principle by extracting and eliminating the requirement for velocity and acceleration sensors from the system. The method reformulates the dynamic model to depend only on position measurements, which are already available from standard encoder sensors. This extraction of the problematic measurement requirements simplifies the device while maintaining prediction accuracy through a restructured modeling approach that uses position-only data with augmented state variables.
3Adaptability or versatility
If Gaussian Process Regression models are used for model-based reinforcement learning, then generalization capability is improved, but computational load increases
Solution Approach 1:
The patent applies segmentation by dividing the computational task into distinct phases: offline training phase where the Gaussian Process model is trained using simulation data, and online execution phase where the trained model is used for prediction. The complex GP training computations are performed offline when computational resources are abundant, while the online phase requires minimal computation for real-time control decisions. This temporal segmentation reduces the immediate computational load during robot operation while maintaining high generalization capability.
Data Source
AI summary
A manipulator learning-control apparatus for controlling a manipulating system that includes an interface configured to receive manipulator state signals of the manipulating system and object state signals with respect to an object to be manipulated by the manipulating system in a workspace, wherein the object state signals are detected by at least one object detector, an output interface configured to transmit initial and updated policy programs to the manipulating system, a memory to store computer-executable programs including a data preprocess program, object state history data, manipulator state history data, a Derivative-Free Semi-parametric Gaussian Process (DF-SPGP) kernel learning program, a Derivative-Free Semi-parametric Gaussian Process (DF-SPGP) model learning program, an update-policy program and an initial policy program, and a processor, in connection with the memory, configured to transmit the initial policy program to the manipulating system for initiating a learning process that operates the manipulator system manipulating the object while a preset period of time.


