Automated Spoken Dialog User Experience Evaluation

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Automated spoken dialog systems face challenges in providing a consistent user experience due to subjective variability in human evaluator scores, which affects the robustness and efficiency of evaluating caller experience (CE) in complex interactions.

Innovation Solution

A system and method that involves selecting multiple evaluators to score interactions, using a machine learning algorithm to predict user experience scores based on features derived from interaction recordings, and an adjudication process to reduce scoring differences, allowing for automated estimation of user experience.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If multiple human evaluators are used to score interactions, then measurement precision of user experience improves, but device complexity and loss of time increase

Engineering Contradiction:
Improveuser experience score accuracyVSAvoidevaluation system complexity
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

An automated evaluation system acts as an intermediary between human evaluators and the final user experience assessment. The system collects scores from multiple human evaluators, applies statistical analysis to reduce variability, and produces a consolidated user experience score, thereby maintaining measurement precision while reducing the complexity and time burden of manual evaluation processes

Inventive Principle:
Principle #24Intermediary (Mediator)

Solution Approach 2:

The patent creates a computational model that replicates the evaluation function of human judges. By training an automated system on examples of human scoring patterns and applying statistical methods to aggregate multiple human scores, the system produces scores that copy the essence of human evaluation without requiring actual human involvement in each scoring instance

Inventive Principle:
Principle #26Copying

2Reliability

If multiple human evaluators are used to score interactions, then reliability of user experience estimates improves, but productivity decreases

Engineering Contradiction:
Improveuser experience estimate robustnessVSAvoidevaluation throughput
Core Design Contradiction:
ReliabilityVSProductivity

Solution Approach 1:

The system performs preliminary actions by collecting and analyzing scores from multiple human evaluators in advance. Statistical methods are applied beforehand to establish reliable baseline scores and variability metrics, so that when actual user experience assessment is needed, the system can quickly produce reliable estimates without requiring repeated manual evaluation of the same interactions

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The evaluation system serves itself by using the aggregated data from initial human evaluations to train and refine its automated scoring algorithms. The system learns from human evaluator patterns and progressively reduces its dependence on manual scoring, thereby maintaining reliability while increasing productivity through automated self-improvement

Inventive Principle:
Principle #25Self-service

3Measurement precision

If subjective variability in human evaluator scores is reduced, then measurement precision improves, but loss of time in adjudication processes increases

Engineering Contradiction:
Improvescore consistencyVSAvoidadjudication time
Core Design Contradiction:
Measurement precisionVSLoss of time

Solution Approach 1:

The system implements feedback mechanisms where adjudication outcomes and score variations are fed back into the statistical model. When score variability exceeds thresholds, the system automatically triggers adjudication processes, and the results of these adjudications are used to refine scoring guidelines and retrain evaluators, thereby reducing future variability and minimizing the need for repeated adjudication

Inventive Principle:
Principle #23Feedback

Solution Approach 2:

The patent dynamically adjusts evaluation parameters such as scoring thresholds, variability tolerances, and adjudication triggers based on accumulated data. By changing these parameters adaptively, the system optimizes the balance between measurement precision and time investment, applying rigorous adjudication only when necessary while maintaining high score consistency through statistical normalization and evaluator calibration

Inventive Principle:
Principle #35Parameter changes

Data Source

PatentUS8520808B2System and method for robust evaluation of the user experience in automated spoken dialog systems
Publication Date: 2013.08.27 VERINT AMERICAS INC
  • US8520808B2 patent drawing
  • US8520808B2 patent drawing
  • US8520808B2 patent drawing

AI summary

A single, subjective numerical rating to evaluate the performance of a telephone-based spoken dialog system. This CE rating is provided by expert human listeners who have knowledge of the design of the dialog system. Different human raters can be trained to achieve a satisfactory level of agreement. Furthermore, a classifier trained on ratings by human experts can reproduce the human ratings with the same degree of consistency. More calls can be given a CE rating than would be possible with limited human resources. More information can be provided about individual calls, e.g., to help decide between two disparate ratings by different human experts.