Cramér-von Mises Divergence for Spoken Dialog Simulation Evaluation

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

There is no accepted method to evaluate the quality of user simulations in spoken dialog systems, which hinders the assessment of machine learning-based dialog systems trained and evaluated on simulated users rather than real users.

Innovation Solution

A novel method using the Cramér-von Mises divergence to measure the similarity between real and simulated dialog scores, providing a domain-specific scoring function to assess the quality of user simulations without assuming the shape of the distribution, and enabling the comparison of user simulations across different domains.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Productivity

If user simulations are used to train and evaluate machine learning-based dialog systems, then the feasibility of training is improved (avoiding the need for thousands of real user dialogs), but the reliability of evaluation deteriorates (no accepted method to assess simulation quality)

Engineering Contradiction:
Improvetraining feasibilityVSAvoidevaluation reliability
Core Design Contradiction:
ProductivityVSReliability

Solution Approach 1:

The patent creates a copy of real user dialog data through user simulations, allowing machine learning training without real user involvement. The simulation generates synthetic dialog data that mimics real user behavior patterns, enabling scalable training while maintaining evaluation capability through comparison with actual user interactions.

Inventive Principle:
Principle #26Copying

Solution Approach 2:

The patent transforms the evaluation approach by changing the parameter being measured from traditional dialog metrics to distributional divergence metrics (KL divergence, Cramér-von Mises divergence). This parameter change enables quantitative assessment of simulation quality by measuring how closely the simulated dialog score distribution matches the real user dialog score distribution.

Inventive Principle:
Principle #35Parameter changes

2Ease of operation

If traditional dialog evaluation metrics are used, then the evaluation process is simple, but the measurement precision deteriorates (cannot capture distributional differences in dialog scores)

Engineering Contradiction:
Improveevaluation simplicityVSAvoiddistribution similarity measurement
Core Design Contradiction:
Ease of operationVSMeasurement precision

Solution Approach 1:

The patent replaces traditional mechanical evaluation methods (manual assessment, simple metric comparison) with statistical distribution analysis. By substituting the evaluation mechanism to use divergence metrics that compare entire score distributions rather than single values, the system achieves both simplicity (automated computation) and precision (capturing distributional characteristics).

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

Data Source

PatentUS8660844B2System and method of evaluating user simulations in a spoken dialog system with a diversion metric
Publication Date: 2014.02.25 AT&T LABS INC
  • US8660844B2 patent drawing
  • US8660844B2 patent drawing
  • US8660844B2 patent drawing

AI summary

Systems, methods and computer-readable media associated with using a divergence metric to evaluate user simulations in a spoken dialog system. The method employs user simulations of a spoken dialog system and includes aggregating a first set of one or more scores from a real user dialog, aggregating a second set of one or more scores from a simulated user dialog associated with a user model, determining a similarity of distributions associated with each of the first set and the second set, wherein the similarity is determined using a divergence metric that does not require any assumptions regarding a shape of the distributions. It is preferable to use a Cramér-von Mises divergence.