Robot Control Learning for Low-Intervention Dynamic Navigation

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing robot control methods in dynamic environments, such as crowded spaces, often require frequent re-planning and intervention behaviors, leading to increased stress on the environment and reduced transport efficiency, as they primarily focus on passive collision avoidance and do not effectively manage interactions with multiple agents.

Innovation Solution

A robot control model learning method that uses reinforcement learning to minimize intervention behaviors by assigning negative rewards to interventions, allowing the robot to autonomously navigate through dynamic environments while reducing the frequency of interactions with its surroundings, thereby optimizing arrival time and collision avoidance.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Reliability

If frequent intervention behaviors are used to navigate through dynamic environments, then the robot can reach the destination, but the transport efficiency deteriorates and environmental stress increases

Engineering Contradiction:
Improveability to reach destinationVSAvoidtransport efficiency
Core Design Contradiction:
ReliabilityVSProductivity

Solution Approach 1:

The patent changes the reward parameter in the reinforcement learning system by assigning negative rewards to intervention behaviors. This parameter change transforms the robot's behavior from frequent interventions to minimal interventions, improving transport efficiency while maintaining the ability to reach destinations through learned navigation strategies

Inventive Principle:
Principle #35Parameter changes

Solution Approach 2:

The patent implements feedback mechanisms where the robot receives negative reward signals when performing intervention behaviors. This feedback loop enables the robot to learn from past interventions and reduce future interventions, optimizing the balance between reaching destinations and maintaining transport efficiency

Inventive Principle:
Principle #23Feedback

2Reliability

If frequent intervention behaviors are used to navigate through dynamic environments, then the robot can avoid collisions, but environmental stress increases

Engineering Contradiction:
Improvecollision avoidance capabilityVSAvoidenvironmental stress
Core Design Contradiction:
ReliabilityVSObject-affected harmful factors

Solution Approach 1:

The patent modifies the reward parameter structure to assign negative rewards to intervention behaviors. This parameter change enables the robot to maintain collision avoidance capability through learned behaviors while significantly reducing environmental stress by minimizing unnecessary interventions in the dynamic environment

Inventive Principle:
Principle #35Parameter changes

Solution Approach 2:

The robot uses reinforcement learning to develop self-service navigation capabilities, learning to navigate through dynamic environments and avoid collisions autonomously without relying on frequent intervention behaviors, thereby reducing environmental stress

Inventive Principle:
Principle #25Self-service

3Reliability

If passive collision avoidance behaviors are used, then the robot can navigate safely, but the number of interventions increases

Engineering Contradiction:
Improvecollision avoidanceVSAvoidnumber of interventions
Core Design Contradiction:
ReliabilityVSLoss of time

Solution Approach 1:

The patent changes the reward parameter assignment to penalize intervention behaviors with negative rewards. This parameter change transforms the robot from passive collision avoidance with frequent interventions to active navigation with minimal interventions, reducing time loss while maintaining safety

Inventive Principle:
Principle #35Parameter changes

Solution Approach 2:

The robot performs preliminary learning through reinforcement learning before actual navigation tasks. This preliminary action enables the robot to pre-learn optimal navigation strategies that minimize interventions, allowing it to execute tasks with fewer interruptions and less time loss

Inventive Principle:
Principle #10Preliminary action

Data Source

PatentEP4060446B1Robot control model learning method, robot control model learning device, robot control model learning program, robot control method, robot control device, robot control program, and robot
Publication Date: 2024.11.13 OMRON CORP
  • EP4060446B1 patent drawingFigure 1
  • EP4060446B1 patent drawingFigure 2~3
  • EP4060446B1 patent drawingFigure 4

AI summary

A robot control model learning device (10) performs, by using state information indicating the state of a robot which autonomously travels to a destination in a dynamic environment as an input, reinforcement learning to obtain a robot control model for selecting and outputting a behavior in accordance with the state of the robot from among a plurality of behaviors including an intervention behavior for intervening in the environment, while using the number of times the intervention behavior has been performed as a minus reward.