Dialog Conformance Scoring for Conversational Quality Assessment
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current speech processing systems lack effective methods to objectively measure and improve the conversational quality of dialogs, making it difficult for developers to assess and enhance user interactions.
Innovation Solution
The system introduces scoring metrics such as productivity, relevance, and naturalness, using deterministic algorithms and machine learning models to evaluate dialog quality, allowing developers to assess and improve conversational quality by selecting appropriate skills and refining user interactions.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If speech processing systems use traditional evaluation methods, then development process is simple, but conversational quality assessment is subjective and inaccurate
Solution Approach 1:
The conversational quality assessment is divided into multiple independent scoring metrics including productivity score, relevance score, and naturalness score. Each metric evaluates a specific aspect of dialog quality separately, allowing for precise measurement of different dimensions without requiring a monolithic complex system.
Solution Approach 2:
The patent introduces scoring components as intermediary elements that bridge the gap between raw dialog data and quality assessment. These components compute individual scores based on specific criteria, then aggregate them into an overall conversational quality metric, enabling objective evaluation without direct subjective judgment.
2Reliability
If developers manually assess dialog quality, then system complexity remains low, but assessment consistency and objectivity deteriorate
Solution Approach 1:
The system implements automated scoring components that provide consistent feedback on dialog quality across different interactions. Each scoring component applies the same criteria uniformly, ensuring that assessments are reliable and reproducible regardless of who or what performs the evaluation, thereby eliminating human subjectivity while maintaining systematic complexity.
3Measurement precision
If the system evaluates all aspects of dialog quality comprehensively, then assessment accuracy improves, but processing time increases
Solution Approach 1:
By segmenting the quality assessment into parallel scoring components (productivity, relevance, naturalness), the system can evaluate multiple aspects simultaneously rather than sequentially. This parallel processing approach maintains comprehensive assessment accuracy while reducing total processing time compared to sequential evaluation methods.
Data Source
AI summary
Techniques for generating a conformance score for a system/user dialog are described. A conformance score may represent a degree to which output data, provided by a skill, conforms to various policies (e.g., the data includes content appropriate for the age of the user, the data does not include profanity, etc.). User input data and system output data, corresponding to a dialog exchange between a user and a skill, ma be determined. A user type associated with the dialog exchange may also be determined. Based on the user type and the system output data, it may be determined that one or more filtering resources are to be assigned to process future data, received from the skill, prior to the future data being presented to a user.


