Reinforcement Learning for Robot Handover Timing Optimization

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current robot systems face challenges in optimizing handover positions and times during cooperative work with operators, requiring repetitive programming and limiting efficiency.

Innovation Solution

An action information learning device that uses reinforcement learning to acquire and adjust state information, calculate rewards, and update value functions to minimize handover time, facilitating optimized action information output for robot control systems.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Productivity

If handover position and timing are optimized through programming, then cooperative work efficiency is improved, but programming time and system complexity increase

Engineering Contradiction:
Improvecooperative work efficiencyVSAvoidprogramming complexity
Core Design Contradiction:
ProductivityVSDevice complexity

Solution Approach 1:

The robot performs self-learning of optimal handover positions and timings through reinforcement learning. The learning unit automatically acquires state information from sensors, calculates rewards based on handover performance, and updates the value function without human intervention, enabling the system to optimize itself autonomously

Inventive Principle:
Principle #25Self-service

Solution Approach 2:

The system implements continuous feedback loops where the learning unit monitors handover state information, calculates reward values based on handover timing and success, and uses this feedback to iteratively update the value function and improve future handover decisions

Inventive Principle:
Principle #23Feedback

2Measurement precision

If handover position is adjusted through trial and error, then optimal position can be found, but time consumption and productivity are reduced

Engineering Contradiction:
Improvehandover position accuracyVSAvoidoptimization speed
Core Design Contradiction:
Measurement precisionVSProductivity

Solution Approach 1:

The system performs preliminary acquisition of state information related to handover positions and uses the value function to pre-calculate optimal actions, avoiding the need for time-consuming trial and error during actual operation

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The patent replaces physical trial-and-error adjustments with an information-based reinforcement learning system that uses sensors, value function calculations, and automated decision-making to determine optimal handover positions

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

3Loss of time

If reinforcement learning is implemented to optimize handover, then handover time is reduced, but system complexity and computational requirements increase

Engineering Contradiction:
Improvehandover timeVSAvoidlearning system complexity
Core Design Contradiction:
Loss of timeVSDevice complexity

Solution Approach 1:

The learning system is segmented into distinct functional units: state information acquisition unit for sensing, learning unit for processing divided into reward calculation and value function update sections, and action information output unit for execution, allowing modular implementation and management of complexity

Inventive Principle:
Principle #1Segmentation

Data Source

PatentUS10730182B2Action information learning device, robot control system and action information learning method
Publication Date: 2020.08.04 FANUC LTD
  • US10730182B2 patent drawing
  • US10730182B2 patent drawing
  • US10730182B2 patent drawing

AI summary

To provide an action information learning device, robot control system and action information learning method for facilitating the performing of cooperative work by an operator with a robot. An action information learning device includes: a state information acquisition unit that acquires a state of a robot; an action information output unit for outputting an action, which is adjustment information for the state; a reward calculation section for acquiring determination information, which is information about a handover time related to handover of a workpiece, and calculating a value of reward in reinforcement learning based on the determination information thus acquired; and a value function update section for updating a value function by way of performing the reinforcement learning based on the value of reward calculated by the reward calculation section, the state and the action.