Machine-Generated Audio Transcript Ranking With SVM Thresholds
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing automatic speech recognition systems lack a consistent and qualitatively meaningful method to evaluate the performance of transcription models, as subjective user evaluations vary and quantitative measures like word-error rates (WERs) do not provide clear categorical assessments.
Innovation Solution
Implementing a support vector machine (SVM) model to translate objective, quantitative accuracy measures into qualitative categories by defining specific threshold boundaries, using both objective and subjective performance indications to categorize transcripts.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Ease of operation
If subjective user evaluations are used to assess transcription quality, then user perspective is captured, but consistency and objectivity are lost
Solution Approach 1:
The patent segments the evaluation process into two distinct components: subjective user evaluations (capturing user perspective) and objective automated metrics (providing consistency). By separating these evaluation methods and combining their results, the system maintains both user relevance and evaluation reliability without requiring one to completely replace the other.
2Reliability
If quantitative error rates are used to measure transcription accuracy, then objectivity is achieved, but qualitative meaningfulness is lost
Solution Approach 1:
The patent transforms the evaluation parameters by introducing weighted scoring mechanisms that convert raw quantitative error rates into qualitatively meaningful categories. By applying different weights to different types of errors and transforming the output into ranked categories, the system maintains objective measurement while recovering qualitative interpretability.
3Measurement precision
If multiple evaluation metrics are combined to assess transcription quality, then comprehensive assessment is achieved, but system complexity increases
Solution Approach 1:
The patent merges multiple evaluation metrics (subjective user ratings, objective error rates, and qualitative assessments) into a unified weighted scoring system. By combining these diverse metrics into a single comprehensive score with defined weightings, the system achieves comprehensive assessment while managing complexity through integration rather than separate parallel processes.
Data Source
AI summary
In general, this disclosure describes techniques for generating and evaluating automatic transcripts of audio recordings containing human speech. In some examples, a computing system is configured to: generate transcripts of a plurality of audio recordings; determine an error rate for each transcript by comparing the transcript to a reference transcript of the audio recording; receive, for each transcript, a subjective ranking selected from a plurality of subjective rank categories; determine, based on the error rates and subjective rankings, objective rank categories defined by error-rate ranges; and assign an objective ranking to a new machine-generated transcript of a new audio recording, based on the objective rank categories and an error rate of the new machine-generated transcript.


