Dialog Apparatus Using Reinforcement Learning for Adaptive Response Selection

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Conventional inquiry dialog systems operate based on static rules predefined by designers, limiting their ability to adapt to diverse user types and goals, and manually creating adaptable rules is challenging due to high development time and cost.

Innovation Solution

A dialog apparatus and method that includes a policy unit scoring and selecting response candidates based on dialog state and policy parameters, and a policy parameter updating unit using a reward function to evaluate behavior and update parameters, allowing the system to adapt to specific circumstances.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Adaptability or versatility

If static rules are used in the policy unit, then the system structure remains simple and easy to implement, but the system cannot adapt to diverse user types and goals

Engineering Contradiction:
Improveadaptability to diverse user types and goalsVSAvoidsystem structure complexity
Core Design Contradiction:
Adaptability or versatilityVSDevice complexity

Solution Approach 1:

The patent transforms the static policy unit into a dynamic reinforcement learning-based policy that can adapt to different user types and goals. The policy unit now includes a policy network that learns optimal policies through interaction with the environment, enabling the system to handle diverse inquiry scenarios without manually reconfiguring rules for each case.

Inventive Principle:
Principle #15Dynamics

Solution Approach 2:

The patent changes the parameters of the policy unit from fixed static rules to learnable parameters in a neural network. The policy network's weights and biases are updated through reinforcement learning based on rewards received from successful dialog outcomes, allowing the system to adapt its behavior to different user types and goals by changing its internal parameters rather than its structure.

Inventive Principle:
Principle #35Parameter changes

2Adaptability or versatility

If manual rule creation is performed to cover various circumstances, then the system can handle diverse scenarios, but development time and cost increase significantly

Engineering Contradiction:
Improvecoverage of various circumstancesVSAvoiddevelopment time
Core Design Contradiction:
Adaptability or versatilityVSLoss of time

Solution Approach 1:

The patent implements self-service by enabling the policy unit to automatically learn and adapt to various circumstances through reinforcement learning. Instead of requiring developers to manually create rules for every scenario, the system autonomously improves its performance by learning from interactions with users and the environment, significantly reducing development time and effort.

Inventive Principle:
Principle #25Self-service

Solution Approach 2:

The patent introduces feedback mechanisms where the policy unit receives reward signals based on the outcome of dialog interactions. This feedback loop allows the system to learn from successful and unsuccessful dialog outcomes, automatically adjusting its policy to handle diverse circumstances without manual intervention, thereby reducing development time while maintaining high adaptability.

Inventive Principle:
Principle #23Feedback

3Adaptability or versatility

If manual rule creation is performed to cover various circumstances, then the system can handle diverse scenarios, but development cost increases significantly

Engineering Contradiction:
Improvecoverage of various circumstancesVSAvoiddevelopment cost
Core Design Contradiction:
Adaptability or versatilityVSEase of manufacture

Solution Approach 1:

The patent implements self-service by enabling the policy unit to automatically learn and adapt to various circumstances through reinforcement learning. Instead of requiring developers to manually create rules for every scenario, the system autonomously improves its performance by learning from interactions with users and the environment, significantly reducing development time and effort.

Inventive Principle:
Principle #25Self-service

Solution Approach 2:

The patent replaces the mechanical process of manual rule creation with an automated reinforcement learning system. The policy network learns optimal policies through trial and error in the dialog environment, substituting the manual intellectual labor of rule creators with an automated learning process that reduces both time and cost while improving adaptability.

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

4Adaptability or versatility

If the policy unit operates based on predefined rules, then the system is easy to implement, but it cannot adapt to new circumstances not anticipated by the designer

Engineering Contradiction:
Improveadaptation to new circumstancesVSAvoidpolicy unit complexity
Core Design Contradiction:
Adaptability or versatilityVSDevice complexity

Solution Approach 1:

The patent transforms the static policy unit into a dynamic reinforcement learning-based policy that can adapt to different user types and goals. The policy unit now includes a policy network that learns optimal policies through interaction with the environment, enabling the system to handle diverse inquiry scenarios without manually reconfiguring rules for each case.

Inventive Principle:
Principle #15Dynamics

Solution Approach 2:

The patent changes the parameters of the policy unit from fixed static rules to learnable parameters in a neural network. The policy network's weights and biases are updated through reinforcement learning based on rewards received from successful dialog outcomes, allowing the system to adapt its behavior to different user types and goals by changing its internal parameters rather than its structure.

Inventive Principle:
Principle #35Parameter changes

Data Source

PatentUS11663413B2Dialog apparatus, dialog system, and computer-readable recording medium
Publication Date: 2023.05.30 NEC CORP
  • US11663413B2 patent drawing
  • US11663413B2 patent drawing
  • US11663413B2 patent drawing

AI summary

A dialog apparatus 100 is an apparatus for responding to a dialog act of a user. The dialog apparatus 100 is provided with: a policy unit 40 configured to set a score to each of response candidates included in a set of response candidates based on the state of a dialog being performed with the user and a policy parameter, and referring to the set scores, to select one of the response candidates as a dialog act of the dialog apparatus 100; and a policy parameter updating unit 60 configured to obtain a reward in the state of the dialog using a reward function that, as the reward, returns an evaluation of a behavior performed in a specific circumstance as a quantitatively represented numeric value, and to update the policy parameter based on the obtained reward.