Robot Preference Learning Under Noisy Human Actions

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing human-robot interaction systems struggle to effectively learn and adapt to human preferences in real-time, particularly in bilateral interactions, due to inherent noise in human actions and variability in human behavior across different task contexts.

Innovation Solution

A system utilizing a Constrained Partially Observable Markov Decision Process (CPOMDP) with Boltzmann rationality and maximum entropy, combined with radial basis functions and hierarchical optimization, to generate robot actions that align with human preferences by sensing noisy human actions, generating features, updating beliefs, and implementing robot actions through actuators while adhering to constraints.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Adaptability or versatility

If the robot uses traditional HRI systems to learn human preferences, then the system structure is simple, but the robot cannot effectively adapt to human preferences in real-time due to noise in human actions and variability in behavior

Engineering Contradiction:
Improveadaptability to human preferencesVSAvoidsystem complexity
Core Design Contradiction:
Adaptability or versatilityVSDevice complexity

Solution Approach 1:

The system segments the HRI problem into distinct modules: observation model (Boltzmann rationality), belief model (maximum entropy), and control model (hierarchical optimization). Each module handles a specific aspect of learning and adapting to human preferences, allowing the complex adaptation task to be broken down into manageable components that can be solved independently and combined for overall system functionality.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The system dynamically changes parameters such as the observation model based on Boltzmann rationality and belief model based on maximum entropy to adapt to varying human behaviors. The hierarchical optimization adjusts constraint weights and priorities in real-time based on observed human actions, enabling the robot to adapt its control parameters to match human preferences without requiring a complete system redesign.

Inventive Principle:
Principle #35Parameter changes

2Reliability

If the robot operates in uncertain and dynamic environments, then the robot can handle real-time HRI, but the noise in human actions makes it difficult to accurately perceive and respond to human preferences

Engineering Contradiction:
Improvereliability of preference perceptionVSAvoidinformation loss from noisy actions
Core Design Contradiction:
ReliabilityVSLoss of information

Solution Approach 1:

The system implements continuous feedback loops where the robot observes human actions, updates its belief model based on maximum entropy principles, and adjusts its behavior accordingly. The observation model uses Boltzmann rationality to interpret noisy human actions, and the hierarchical optimization provides feedback by adjusting constraint satisfaction based on observed preferences, creating a closed-loop system that progressively improves its understanding of human intentions despite environmental noise.

Inventive Principle:
Principle #23Feedback

Solution Approach 2:

The system performs preliminary actions by pre-establishing constraint hierarchies and observation models before actual HRI occurs. The belief model is initialized with prior knowledge about human preferences, and the hierarchical optimization framework is pre-configured with constraint weights. This preliminary setup enables the robot to quickly adapt to new situations without needing to learn everything from scratch, reducing the impact of noise in real-time observations.

Inventive Principle:
Principle #10Preliminary action

3Reliability

If the robot implements comprehensive constraint optimization for safety, then safety and efficiency are improved, but the computational complexity and processing time increase

Engineering Contradiction:
Improvesafety in HRIVSAvoidcomputational processing time
Core Design Contradiction:
ReliabilityVSLoss of time

Solution Approach 1:

The constraint optimization is segmented into hierarchical levels with different priorities. Critical safety constraints (such as collision avoidance) are handled at higher levels with stricter enforcement, while less critical constraints (such as operational efficiency) are handled at lower levels. This segmentation allows the system to quickly satisfy safety requirements without needing to optimize all constraints simultaneously, reducing computational time while maintaining safety.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The system applies partial optimization by focusing computational resources on the most critical constraints rather than equally optimizing all constraints. The hierarchical optimization framework allows the robot to satisfy essential safety constraints thoroughly while applying lighter optimization to secondary constraints. This partial action approach ensures safety is not compromised while reducing overall computational processing time.

Inventive Principle:
Principle #16Partial or excessive action

Data Source

PatentUS20250345931A1Learning perceived preferences in human-robot interactions (HRI)
Publication Date: 2025.11.13 HONDA MOTOR CO LTD
  • US20250345931A1 patent drawing
  • US20250345931A1 patent drawing
  • US20250345931A1 patent drawing

AI summary

According to one aspect, learning perceived preferences in human-robot interactions (HRI) may include sensing a noisy action from a human associated with a human-robot interaction (HRI) with a robot, generating a feature associated with the human based on the noisy action and an observation model, generating a belief based on the feature and a belief model, generating a robot action based on a reference trajectory, the belief, and one or more constraints, and implementing the robot action for the HRI via a robot appendage of the robot and an actuator.