Speech Forced Alignment Model Evaluation via Phoneme Timestamp Accuracy

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current methods for evaluating speech forced alignment models are time-consuming, labor-intensive, and subject to subjective influences, resulting in low accuracy and high costs.

Innovation Solution

A method and apparatus for evaluating speech forced alignment models by acquiring phoneme sequences and predicted start/end times, calculating time accuracy scores based on reference times, and aggregating these scores to assess model accuracy, thereby simplifying the evaluation process and reducing labor and time costs.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If manual evaluation methods are used for speech forced alignment models, then evaluation can be performed, but the process is time-consuming and labor-intensive

Engineering Contradiction:
Improveevaluation accuracyVSAvoidevaluation time
Core Design Contradiction:
Measurement precisionVSLoss of time

Solution Approach 1:

The patent replaces manual mechanical evaluation processes with an automated computational system. The evaluation apparatus automatically calculates time accuracy scores by comparing predicted phoneme timestamps against reference timestamps, eliminating the need for manual retesting and subjective human evaluation, thus resolving the contradiction between evaluation accuracy and time consumption.

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

2Measurement precision

If manual retesting and subsequent speech synthesis are performed for evaluation, then accuracy can be verified, but labor costs and time costs increase

Engineering Contradiction:
Improvemodel accuracyVSAvoidevaluation cost
Core Design Contradiction:
Measurement precisionVSEase of manufacture

Solution Approach 1:

The patent extracts the essential evaluation metric (time accuracy score) directly from the model's timestamp predictions without requiring full manual retesting or speech synthesis. By taking out only the critical comparison between predicted and reference timestamps, the system achieves accurate model evaluation while eliminating unnecessary labor-intensive steps, thus reducing evaluation costs.

Inventive Principle:
Principle #2Taking out (Extraction)

3Extent of automation

If speech forced alignment technology is used to obtain phoneme time points, then automation is achieved, but the accuracy of phoneme time points is not high

Engineering Contradiction:
Improveautomation levelVSAvoidphoneme time point accuracy
Core Design Contradiction:
Extent of automationVSMeasurement precision

Solution Approach 1:

The patent introduces a feedback mechanism where the evaluation apparatus compares the automated model predictions against reference phoneme timestamps and calculates time accuracy scores. This feedback loop enables identification of accuracy issues in the forced alignment model, allowing for targeted improvements while maintaining automation, thus resolving the contradiction between automation level and prediction accuracy.

Inventive Principle:
Principle #23Feedback

Data Source

PatentUS11749257B2Method for evaluating a speech forced alignment model, electronic device, and storage medium
Publication Date: 2023.09.05 BEIJING CENTURY TAL EDUCATION TECH CO LTD
  • US11749257B2 patent drawing
  • US11749257B2 patent drawing
  • US11749257B2 patent drawing

AI summary

A method for evaluating a speech forced alignment model, an electronic device, and a storage medium are provided. The method includes: according to each audio segment in a test set and a text corresponding to each audio segment, acquiring, by using a speech forced alignment model to be evaluated, a phoneme sequence corresponding to each audio segment and a predicted start time and a predicted end time of each phoneme in the phoneme sequence; for each phoneme, obtaining a time accuracy score of the phoneme according to the predicted start time and the predicted end time of the phoneme and a predetermined reference start time and a predetermined reference end time of the phoneme; and determining a time accuracy score of said speech forced alignment model according to the time accuracy score of each phoneme.