Robot Control Learning to Reduce Intervention in Dynamic Navigation
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing robot path planning techniques struggle in dynamic environments, particularly in crowded spaces, leading to frequent re-planning, stress on the environment, and inefficient transport due to high intervention frequencies.
Innovation Solution
A robot control model learning method that uses reinforcement learning to minimize intervention behaviors by optimizing state value functions and selecting behaviors that reduce collision avoidance and intervention frequency, using negative rewards for intervention actions.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Reliability
If intervention behaviors are frequently executed to avoid collisions in dynamic environments, then collision avoidance capability is improved, but stress on the environment and transport efficiency deteriorate
Solution Approach 1:
The patent applies parameter changes by modifying the reward function parameters in the reinforcement learning system. Specifically, it assigns negative rewards for intervention behaviors and positive rewards for maintaining smooth navigation, thereby changing the behavioral parameters of the robot to reduce unnecessary interventions while maintaining collision avoidance capability
Solution Approach 2:
The patent implements feedback mechanisms by continuously monitoring the robot's state, environment conditions, and intervention frequency. The reinforcement learning system uses this feedback to adjust the policy, learning to distinguish between situations requiring intervention and those where passive avoidance is sufficient, thereby reducing environmental stress
2Ease of operation
If simple intervention policies are implemented for easy execution, then ease of operation is improved, but intervention frequency increases causing stress on the environment
Solution Approach 1:
The patent applies dynamics by transitioning from static, pre-programmed intervention policies to dynamic, adaptive policies learned through reinforcement learning. The system dynamically adjusts intervention behavior based on real-time environmental conditions and learned patterns, maintaining ease of implementation while reducing unnecessary interventions to improve transport efficiency
Solution Approach 2:
The robot performs self-service by autonomously learning the optimal intervention policy through reinforcement learning without requiring external programming or manual adjustment. The system serves itself by automatically optimizing its behavior to balance ease of operation with reduced intervention frequency, thereby improving transport efficiency
3Adaptability or versatility
If re-planning is carried out frequently in congested environments, then path adaptability is improved, but robot stoppage and loss of time increase
Solution Approach 1:
The patent applies preliminary action by pre-training the reinforcement learning model offline to learn optimal navigation strategies for various congested environment scenarios. This preliminary learning enables the robot to make rapid, informed decisions during actual operation without requiring frequent time-consuming re-planning, thus maintaining path adaptability while reducing stoppage time
Solution Approach 2:
The patent substitutes the traditional mechanical re-planning system with an intelligent, learning-based decision-making system. Instead of relying on computationally intensive path re-calculation in congested environments, the robot uses the learned policy to quickly adapt its behavior, replacing the mechanical re-planning process with a more efficient neural network-based approach that reduces loss of time
Data Source
AI summary
A robot control model learning device (10) performs, by using state information indicating the state of a robot which autonomously travels to a destination in a dynamic environment as an input, reinforcement learning to obtain a robot control model for selecting and outputting a behavior in accordance with the state of the robot from among a plurality of behaviors including an intervention behavior for intervening in the environment, while using the number of times the intervention behavior has been performed as a minus reward.


