Cramér-von Mises Divergence for Spoken Dialog Simulation Evaluation
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
There is no accepted method to evaluate the quality of user simulations in spoken dialog systems, which hinders the assessment of machine learning-based dialog systems trained and evaluated on simulated users rather than real users.
Innovation Solution
A novel method using the Cramér-von Mises divergence to measure the similarity between real and simulated dialog scores, providing a domain-specific scoring function to assess the quality of user simulations without assuming the shape of the distribution, and enabling the comparison of user simulations across different domains.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Productivity
If user simulations are used to train and evaluate machine learning-based dialog systems, then the feasibility of training is improved (avoiding the need for thousands of real user dialogs), but the reliability of evaluation deteriorates (no accepted method to assess simulation quality)
Solution Approach 1:
The patent creates a copy of real user dialog data through user simulations, allowing machine learning training without real user involvement. The simulation generates synthetic dialog data that mimics real user behavior patterns, enabling scalable training while maintaining evaluation capability through comparison with actual user interactions.
Solution Approach 2:
The patent transforms the evaluation approach by changing the parameter being measured from traditional dialog metrics to distributional divergence metrics (KL divergence, Cramér-von Mises divergence). This parameter change enables quantitative assessment of simulation quality by measuring how closely the simulated dialog score distribution matches the real user dialog score distribution.
2Ease of operation
If traditional dialog evaluation metrics are used, then the evaluation process is simple, but the measurement precision deteriorates (cannot capture distributional differences in dialog scores)
Solution Approach 1:
The patent replaces traditional mechanical evaluation methods (manual assessment, simple metric comparison) with statistical distribution analysis. By substituting the evaluation mechanism to use divergence metrics that compare entire score distributions rather than single values, the system achieves both simplicity (automated computation) and precision (capturing distributional characteristics).
Data Source
AI summary
Systems, methods and computer-readable media associated with using a divergence metric to evaluate user simulations in a spoken dialog system. The method employs user simulations of a spoken dialog system and includes aggregating a first set of one or more scores from a real user dialog, aggregating a second set of one or more scores from a simulated user dialog associated with a user model, determining a similarity of distributions associated with each of the first set and the second set, wherein the similarity is determined using a divergence metric that does not require any assumptions regarding a shape of the distributions. It is preferable to use a Cramér-von Mises divergence.


