Reinforcement Learning Apparatus Using Virtual External Force for Obstacle Avoidance

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Conventional reinforcement learning techniques face challenges in setting a balanced reward function for complex motion trajectories, particularly when obstacle avoidance is required, leading to inefficient learning and potential collisions or failure to move due to the trade-off between reward elements.

Innovation Solution

A reinforcement learning apparatus that utilizes a virtual external force to simplify the reward function, allowing for quick and stable robot motor learning by separating the reinforcement learner and virtual external force generator, and incorporating a virtual external force approximator to adapt and reuse learning results, thereby reducing the need for extensive relearning when obstacles change.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Adaptability or versatility

If a complex reward function with multiple terms is used to express requirements such as obstacle avoidance and reaching movement, then the learning can address multiple tasks, but the trade-off between terms impedes learning speed and stability

Engineering Contradiction:
Improveability to handle multiple tasksVSAvoidlearning speed
Core Design Contradiction:
Adaptability or versatilityVSProductivity

Solution Approach 1:

The patent segments the control inputs into two distinct components: reinforcement learning outputs and virtual external force outputs. The virtual external force specifically handles obstacle avoidance requirements, while the reinforcement learning focuses on task-specific behaviors. This segmentation allows each component to specialize in particular aspects of control, eliminating the trade-off problem that occurs when a single complex reward function tries to balance multiple competing requirements.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The virtual external force acts as an intermediary that mediates between the robot and obstacles. Instead of encoding obstacle avoidance directly into the reward function, the system introduces a virtual external force field that automatically pushes the robot away from obstacles. This intermediary mechanism handles the obstacle avoidance requirement separately, allowing the reinforcement learning to focus on task completion without being impeded by trade-offs.

Inventive Principle:
Principle #24Intermediary (Mediator)

2Productivity

If the reward function is simplified to improve learning speed, then learning becomes faster and more stable, but the ability to handle complex requirements such as obstacle avoidance is reduced

Engineering Contradiction:
Improvelearning speedVSAvoidability to handle complex requirements
Core Design Contradiction:
ProductivityVSAdaptability or versatility

Solution Approach 1:

The patent merges two separate control systems: a simplified reinforcement learning system and a virtual external force system. The reinforcement learning component uses a simple reward function focused on task completion, while the virtual external force component handles complex environmental constraints like obstacle avoidance. By combining these two systems, the patent achieves both fast, stable learning (from the simple reward function) and the ability to handle complex requirements (from the virtual external force).

Inventive Principle:
Principle #5Merging (Combining)

3Manufacturing precision

If empirical adjustment of reward function elements is performed to balance trade-offs, then the reward function can be tuned for specific tasks, but the advantage of autonomous reinforcement learning is compromised

Engineering Contradiction:
Improvereward function tuning precisionVSAvoidautonomous learning capability
Core Design Contradiction:
Manufacturing precisionVSExtent of automation

Solution Approach 1:

The virtual external force system provides self-service by automatically generating appropriate forces based on the robot's state and environment, without requiring manual tuning of reward parameters. The system autonomously determines when and how strongly to apply virtual external forces to prevent collisions, eliminating the need for empirical adjustment of reward function elements while maintaining full autonomous learning capability.

Inventive Principle:
Principle #25Self-service

Data Source

PatentUS8886357B2Reinforcement learning apparatus, control apparatus, and reinforcement learning method
Publication Date: 2014.11.11 ATR ADVANCED TELECOMM RES INST INT
  • US8886357B2 patent drawing
  • US8886357B2 patent drawing
  • US8886357B2 patent drawing

AI summary

It is possible to perform robot motor learning in a quick and stable manner using a reinforcement learning apparatus including: a first-type environment parameter obtaining unit that obtains a value of one or more first-type environment parameters; a control parameter value calculation unit that calculates a value of one or more control parameters maximizing a reward by using the value of the one or more first-type environment parameters; a control parameter value output unit that outputs the value of the one or more control parameters to the control object; a second-type environment parameter obtaining unit that obtains a value of one or more second-type environment parameters; a virtual external force calculation unit that calculates the virtual external force by using the value of the one or more second-type environment parameters; and a virtual external force output unit that outputs the virtual external force to the control object.