Automatic Speech Assessment System for Intelligibility Scoring
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Conventional methods for assessing speech intelligibility are cumbersome and inaccurate, often relying on subjective human evaluations or machine learning techniques that require cumbersome manual labeling of data, leading to inconsistent and unreliable results.
Innovation Solution
An automatic speech assessment system generates an intelligibility score based on an N-best output from an ASR module, calculating conditional intelligibility values and adjusting scores with confidence and pronunciation scores to provide an objective measure of speech clarity.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If human evaluators are used to judge speech intelligibility, then subjective assessment can be performed, but the results are highly subjective and inconsistent across different evaluators
Solution Approach 1:
The patent replaces the mechanical human evaluation process with an automated speech assessment system that uses ASR technology to objectively measure speech intelligibility. The system processes acoustic signals through computational algorithms rather than human perception, eliminating subjective variability while maintaining measurement precision through standardized automated procedures.
Solution Approach 2:
The system enables self-assessment of speech intelligibility by automatically processing and evaluating speech samples without requiring external human evaluators. The automated system performs the assessment function that previously required human judgment, providing consistent and reproducible results through algorithmic processing.
2Extent of automation
If machine learning techniques with manual labeling are used, then automated assessment can be achieved, but the process becomes cumbersome and requires extensive manual data labeling
Solution Approach 1:
The patent extracts and eliminates the cumbersome manual labeling step from the machine learning process. Instead of requiring manually labeled training data, the system uses ASR-generated transcripts and automatic alignment algorithms to directly assess intelligibility, removing the bottleneck of manual data preparation while maintaining automated assessment capabilities.
Solution Approach 2:
The system creates synthetic training data by copying and adapting ASR output and alignment results instead of requiring original manual annotations. This allows the system to generate sufficient training examples through automated processes, reducing dependency on extensive manual labeling while maintaining the benefits of machine learning.
3Ease of operation
If only 1-best ASR output is used for intelligibility scoring, then the process is simple, but the accuracy is reduced due to inability to account for alternative interpretations
Solution Approach 1:
The patent segments the intelligibility assessment process into multiple components: generating multiple ASR hypotheses (N-best output), creating alternative alignments for each hypothesis, and computing intelligibility scores for each segment. This segmentation allows the system to consider multiple interpretations while maintaining a structured, manageable process that balances simplicity with accuracy.
Solution Approach 2:
The system performs more alignment operations than strictly necessary by generating alignments for multiple ASR hypotheses rather than just the single best match. This excessive action of creating multiple alternative alignments ensures that no plausible interpretation is missed, improving measurement precision while keeping the overall process computationally feasible.
Data Source
AI summary
In a method for efficiently and accurately measuring the intelligibility of speech, a user may utter a sample text, and an automatic speech assessment (ASA) system may receive an acoustic signal encoding the utterance. An automatic speech recognition (ASR) module may generate an N-best output corresponding to the utterance and generate an intelligibility score representing the intelligibility of the utterance based on the N-best output and the sample text. Generating the intelligibility score may involve (1) calculating conditional intelligibility value(s) for the N recognition result(s), and (2) determining the intelligibility score based on the conditional intelligibility value of the most intelligible recognition result. Optionally, the process of generating the intelligibility score may involve adjusting the intelligibility score to account for environmental information (e.g., a pronunciation score for the user's speech and/or a confidence score assigned to the 1-best recognition result). N may be greater than or equal to 2.


