Inverse Reinforcement Learning with Differentiable Constraint Distributions

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current inverse reinforcement learning methods face challenges in simultaneously learning an appropriate reward function and constraint conditions from human decision-making history data, especially when constraint conditions are implicitly set or contain noise, leading to difficulties in estimating objective functions and constraint conditions accurately.

Innovation Solution

A learning device and method that performs inverse reinforcement learning using trajectory data, where a differentiable function indicates the distribution of constraint conditions, allowing simultaneous learning of reward functions and constraint conditions through a probabilistic model defined by the maximum entropy principle, enabling the estimation of both parameters from demonstration data.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If constraint conditions are implicitly set or contain noise in trajectory data, then inverse reinforcement learning cannot accurately estimate objective functions and constraint conditions, but requiring constraint conditions to be known in advance limits adaptability

Engineering Contradiction:
Improveestimation accuracy of objective function and constraint conditionsVSAvoidability to handle noisy or varying demonstration data
Core Design Contradiction:
Measurement precisionVSAdaptability or versatility

Solution Approach 1:

The patent transforms constraint conditions from fixed, known parameters into learnable parameters with probability distributions. By representing constraint conditions as probabilistic parameters that can be inferred from trajectory data, the system adapts to noisy and varying data while maintaining estimation accuracy. The constraint conditions are modeled as distributions rather than deterministic values, allowing the system to handle uncertainty and noise in the demonstration data.

Inventive Principle:
Principle #35Parameter changes

2Productivity

If multiple candidate constraint conditions are prepared in advance, then learning can be performed efficiently, but the method cannot adapt when candidate conditions deviate from actual assumptions

Engineering Contradiction:
Improvelearning efficiencyVSAvoidadaptability when candidate conditions are incorrect
Core Design Contradiction:
ProductivityVSAdaptability or versatility

Solution Approach 1:

The patent makes the constraint conditions dynamic by representing them as probability distributions with learnable parameters rather than static, pre-defined candidates. The system dynamically infers the appropriate constraint conditions from the trajectory data, allowing adaptation to the actual task requirements without being limited to pre-specified candidates. This dynamic approach maintains learning efficiency while significantly improving adaptability.

Inventive Principle:
Principle #15Dynamics

3Ease of operation

If all demonstration data are assumed to be mathematically optimal solutions, then constraint conditions can be learned, but the method fails when data contains noise, non-stationarity, or failure data

Engineering Contradiction:
Improvesimplicity of learning methodVSAvoidrobustness to noisy, non-stationary, or failure data
Core Design Contradiction:
Ease of operationVSReliability

Solution Approach 1:

The patent converts the harmful effect of noise and variations in demonstration data into a beneficial feature by modeling constraint conditions as probability distributions. Rather than treating noise as an error to be eliminated, the system uses the variations in data to infer the underlying probability distributions of constraint conditions. This approach transforms data imperfections into useful information for learning robust constraint conditions that generalize well.

Inventive Principle:
Principle #22Blessing in disguise (Convert harm into benefit)

Data Source

PatentUS20240202504A1Learning device, learning method, and learning program
Publication Date: 2024.06.20 NEC CORP
  • US20240202504A1 patent drawing
  • US20240202504A1 patent drawing
  • US20240202504A1 patent drawing

AI summary

The input means 81 accepts input of trajectory data indicating the subject's decision-making history. The learning means 82 performs inverse reinforcement learning using the trajectory data. The output means 83 outputs a reward function and a constraint condition derived by inverse reinforcement learning. Here, the learning means 82 performs inverse reinforcement learning based on distribution of the trajectory data calculated using a differentiable function that indicates distribution of the constraint condition.