Hierarchical Mixtures of Experts Learning Accuracy

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current methods for learning hierarchical mixtures of experts using inverse reinforcement learning lack sufficient accuracy, particularly when combined with model-free approaches and importance sampling, which can lead to suboptimal estimation of reward functions.

Innovation Solution

A learning device and method that incorporates an EM algorithm for initial learning, switching to factorized asymptotic Bayesian inference when predetermined conditions are met, to enhance the estimation accuracy of hierarchical mixtures of experts by inverse reinforcement learning.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If inverse reinforcement learning is combined with hierarchical mixtures of experts using traditional methods, then the model can handle complex intentions, but the estimation accuracy is insufficient

Engineering Contradiction:
Improveestimation accuracyVSAvoidmodel complexity
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The patent segments the learning process into two distinct phases: EM algorithm for initial parameter estimation and factorized asymptotic Bayesian inference for refinement. This segmentation allows each method to focus on specific aspects of the learning task, improving overall estimation accuracy while managing complexity through structured progression

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent applies preliminary action by using the EM algorithm to obtain initial parameter estimates before applying factorized asymptotic Bayesian inference. This preliminary estimation provides a solid foundation for the subsequent refinement process, ensuring that the complex Bayesian inference starts from a reasonable initial state rather than random parameters

Inventive Principle:
Principle #10Preliminary action

2Adaptability or versatility

If multiple intentions are modeled in the reward function, then the model can represent complex expert behavior, but the reward function becomes difficult to interpret

Engineering Contradiction:
Improveintention representationVSAvoidinterpretability
Core Design Contradiction:
Adaptability or versatilityVSEase of operation

Solution Approach 1:

The hierarchical mixtures of experts structure segments the reward function into multiple expert components, each representing a simple intention. The gate function further segments the selection process, choosing which expert to apply based on the current state. This segmentation maintains interpretability by keeping individual experts simple while achieving versatility through their combination

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

Different parts of the reward function (different experts) have different local qualities or characteristics. Each expert is designed to handle specific types of situations with simple, interpretable reward structures. The gate function dynamically selects which local quality to apply based on the current state, maintaining both versatility and interpretability

Inventive Principle:
Principle #3Local quality

3Adaptability or versatility

If model-free inverse reinforcement learning with importance sampling is used, then the method can be applied broadly, but the estimation accuracy deteriorates

Engineering Contradiction:
Improvemethod applicabilityVSAvoidreward function estimation accuracy
Core Design Contradiction:
Adaptability or versatilityVSMeasurement precision

Solution Approach 1:

The patent changes the parameters and assumptions of the learning method by transitioning from pure model-free importance sampling to factorized asymptotic Bayesian inference. This parameter change involves incorporating asymptotic assumptions and factorized probability structures that significantly improve estimation accuracy while maintaining the broad applicability of the inverse reinforcement learning framework

Inventive Principle:
Principle #35Parameter changes

Data Source

PatentUS20230040914A1Learning device, learning method, and learning program
Publication Date: 2023.02.09 NEC CORP
  • US20230040914A1 patent drawing
  • US20230040914A1 patent drawing
  • US20230040914A1 patent drawing

AI summary

An input unit 81 receives input of a decision-making history of a subject. A learning unit 82 learns hierarchical mixtures of experts by inverse reinforcement learning based on the decision-making history. An output unit 83 outputs the learned hierarchical mixtures of experts. The learning unit 82 learns the hierarchical mixtures of experts using an EM algorithm, and when a learning result using the EM algorithm satisfies a predetermined condition, learns the hierarchical mixtures of experts by factorized asymptotic Bayesian inference.