Robot Control Learning for Low-Intervention Dynamic Navigation
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing robot control methods in dynamic environments, such as crowded spaces, often require frequent re-planning and intervention behaviors, leading to increased stress on the environment and reduced transport efficiency, as they primarily focus on passive collision avoidance and do not effectively manage interactions with multiple agents.
Innovation Solution
A robot control model learning method that uses reinforcement learning to minimize intervention behaviors by assigning negative rewards to interventions, allowing the robot to autonomously navigate through dynamic environments while reducing the frequency of interactions with its surroundings, thereby optimizing arrival time and collision avoidance.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Reliability
If frequent intervention behaviors are used to navigate through dynamic environments, then the robot can reach the destination, but the transport efficiency deteriorates and environmental stress increases
Solution Approach 1:
The patent changes the reward parameter in the reinforcement learning system by assigning negative rewards to intervention behaviors. This parameter change transforms the robot's behavior from frequent interventions to minimal interventions, improving transport efficiency while maintaining the ability to reach destinations through learned navigation strategies
Solution Approach 2:
The patent implements feedback mechanisms where the robot receives negative reward signals when performing intervention behaviors. This feedback loop enables the robot to learn from past interventions and reduce future interventions, optimizing the balance between reaching destinations and maintaining transport efficiency
2Reliability
If frequent intervention behaviors are used to navigate through dynamic environments, then the robot can avoid collisions, but environmental stress increases
Solution Approach 1:
The patent modifies the reward parameter structure to assign negative rewards to intervention behaviors. This parameter change enables the robot to maintain collision avoidance capability through learned behaviors while significantly reducing environmental stress by minimizing unnecessary interventions in the dynamic environment
Solution Approach 2:
The robot uses reinforcement learning to develop self-service navigation capabilities, learning to navigate through dynamic environments and avoid collisions autonomously without relying on frequent intervention behaviors, thereby reducing environmental stress
3Reliability
If passive collision avoidance behaviors are used, then the robot can navigate safely, but the number of interventions increases
Solution Approach 1:
The patent changes the reward parameter assignment to penalize intervention behaviors with negative rewards. This parameter change transforms the robot from passive collision avoidance with frequent interventions to active navigation with minimal interventions, reducing time loss while maintaining safety
Solution Approach 2:
The robot performs preliminary learning through reinforcement learning before actual navigation tasks. This preliminary action enables the robot to pre-learn optimal navigation strategies that minimize interventions, allowing it to execute tasks with fewer interruptions and less time loss
Data Source
Figure 1
Figure 2~3
Figure 4
AI summary
A robot control model learning device (10) performs, by using state information indicating the state of a robot which autonomously travels to a destination in a dynamic environment as an input, reinforcement learning to obtain a robot control model for selecting and outputting a behavior in accordance with the state of the robot from among a plurality of behaviors including an intervention behavior for intervening in the environment, while using the number of times the intervention behavior has been performed as a minus reward.