Machine-Generated Audio Transcript Ranking With SVM Thresholds

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing automatic speech recognition systems lack a consistent and qualitatively meaningful method to evaluate the performance of transcription models, as subjective user evaluations vary and quantitative measures like word-error rates (WERs) do not provide clear categorical assessments.

Innovation Solution

Implementing a support vector machine (SVM) model to translate objective, quantitative accuracy measures into qualitative categories by defining specific threshold boundaries, using both objective and subjective performance indications to categorize transcripts.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Ease of operation

If subjective user evaluations are used to assess transcription quality, then user perspective is captured, but consistency and objectivity are lost

Engineering Contradiction:
Improveuser perspective captureVSAvoidevaluation consistency
Core Design Contradiction:
Ease of operationVSReliability

Solution Approach 1:

The patent segments the evaluation process into two distinct components: subjective user evaluations (capturing user perspective) and objective automated metrics (providing consistency). By separating these evaluation methods and combining their results, the system maintains both user relevance and evaluation reliability without requiring one to completely replace the other.

Inventive Principle:
Principle #1Segmentation

2Reliability

If quantitative error rates are used to measure transcription accuracy, then objectivity is achieved, but qualitative meaningfulness is lost

Engineering Contradiction:
Improveobjective measurementVSAvoidqualitative meaning
Core Design Contradiction:
ReliabilityVSLoss of information

Solution Approach 1:

The patent transforms the evaluation parameters by introducing weighted scoring mechanisms that convert raw quantitative error rates into qualitatively meaningful categories. By applying different weights to different types of errors and transforming the output into ranked categories, the system maintains objective measurement while recovering qualitative interpretability.

Inventive Principle:
Principle #35Parameter changes

3Measurement precision

If multiple evaluation metrics are combined to assess transcription quality, then comprehensive assessment is achieved, but system complexity increases

Engineering Contradiction:
Improvecomprehensive assessmentVSAvoidevaluation system complexity
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The patent merges multiple evaluation metrics (subjective user ratings, objective error rates, and qualitative assessments) into a unified weighted scoring system. By combining these diverse metrics into a single comprehensive score with defined weightings, the system achieves comprehensive assessment while managing complexity through integration rather than separate parallel processes.

Inventive Principle:
Principle #5Merging (Combining)

Data Source

PatentUS20250273215A1Categorizing audio transcriptions
Publication Date: 2025.08.28 WELLS FARGO BANK NA
  • US20250273215A1 patent drawing
  • US20250273215A1 patent drawing
  • US20250273215A1 patent drawing

AI summary

In general, this disclosure describes techniques for generating and evaluating automatic transcripts of audio recordings containing human speech. In some examples, a computing system is configured to: generate transcripts of a plurality of audio recordings; determine an error rate for each transcript by comparing the transcript to a reference transcript of the audio recording; receive, for each transcript, a subjective ranking selected from a plurality of subjective rank categories; determine, based on the error rates and subjective rankings, objective rank categories defined by error-rate ranges; and assign an objective ranking to a new machine-generated transcript of a new audio recording, based on the objective rank categories and an error rate of the new machine-generated transcript.