Robot Motion Planning Using Action-Value Force Control
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Robots equipped with force control and position control struggle with slower motion times compared to those using only position control, necessitating a solution to shorten force-controlled motion times while ensuring safety during operations like packing items without damaging them.
Innovation Solution
A robot motion planning device that utilizes an action-value function and measurement information to determine a target position and calculate a difference, allowing for optimized force control parameters to shorten motion times while ensuring safety by planning a force-controlled motion that avoids collisions and damage.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Reliability
If force control is used to ensure safety during robot operations, then reliability is improved, but motion time increases
Solution Approach 1:
The robot motion planning device pre-calculates optimal motion trajectories and force control parameters before executing the motion. By using reinforcement learning to learn action-value functions in advance and planning force-controlled motions that avoid collisions beforehand, the system ensures safety without real-time computation delays during actual motion execution.
Solution Approach 2:
The system dynamically adjusts motion parameters based on learned patterns. The reinforcement learning agent adapts the action-value function according to task requirements and environmental conditions, optimizing the balance between motion speed and safety. The force control parameters are dynamically modified during motion execution based on pre-planned trajectories.
2Productivity
If force control parameters are optimized to shorten motion time, then productivity is improved, but risk of collision or damage increases
Solution Approach 1:
The system uses reinforcement learning with reward functions that provide feedback on motion performance and safety. The action-value function is updated based on rewards that penalize collisions and damages while rewarding efficient motion. This feedback mechanism learns optimal force control parameters that balance speed and safety over multiple training iterations.
Solution Approach 2:
Safety constraints and collision avoidance rules are incorporated into the motion planning before execution. The system pre-identifies safe motion paths and force limits that prevent collisions and damages, allowing the robot to move quickly within these pre-established safety boundaries without real-time risk assessment delays.
3Adaptability or versatility
If reinforcement learning is used to learn action-value functions, then adaptability is improved, but device complexity increases
Solution Approach 1:
The system uses simulation environments to copy and practice robot operations before deploying learned policies to the actual robot. The reinforcement learning agent learns action-value functions in a virtual environment that mimics real-world physics and constraints, then transfers this learned knowledge to the physical robot, reducing the computational burden on the actual control system.
Solution Approach 2:
The patent replaces complex real-time mechanical control computations with pre-learned action-value functions from reinforcement learning. Instead of calculating optimal control parameters in real-time through complex mechanical models, the system uses the learned function to directly determine appropriate actions based on current state observations, simplifying the control architecture.
Data Source
AI summary
According to one embodiment, a robot motion planning device includes processing circuitry. The processing circuitry receives observation information obtained by observing at least part of a movable range of a robot. The processing circuitry determines, in a case where first observation information is received, a target position to which the robot is to make a motion, using an action-value function and the first observation information. The processing circuitry receives measurement information obtained by measuring a state of the robot, calculates a difference corresponding to the first observation information, using the measurement information, and determines a motion plan of a force-controlled motion of the robot, based on the target position and the difference.


