Text-to-Speech Evaluation Using Machine Learning Acoustic Features

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current text-to-speech (TTS) performance evaluation methods are time-consuming, labor-intensive, and subjective, relying on human evaluators for Mean Opinion Score (MOS) assessments, which lack generalization and are affected by evaluator biases, while objective methods struggle with suitable natural speech references and extracting high-layer characteristics.

Innovation Solution

A system and method using machine learning models, such as Support Vector Machines (SVM) and Deep Neural Networks (DNN), to evaluate TTS systems by extracting relevant acoustic features from speech samples and training models with corresponding scores, applying sub-space decomposition methods like Linear Discriminant Analysis (LDA) to select influential features.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If human evaluators are used for MOS assessment, then evaluation can capture subjective quality aspects, but the process becomes time-consuming and labor-intensive

Engineering Contradiction:
Improvequality assessment accuracyVSAvoidevaluation time
Core Design Contradiction:
Measurement precisionVSLoss of time

Solution Approach 1:

The patent replaces the mechanical human evaluation system with an automated machine learning-based evaluation system. The machine learning model processes acoustic features extracted from speech samples to automatically predict quality scores, eliminating the need for manual listener assessments while maintaining evaluation accuracy.

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

Solution Approach 2:

The patent creates a computational model that copies and replicates the evaluation capabilities of human listeners. The machine learning model learns from training data consisting of acoustic features and corresponding quality scores, enabling it to replicate human evaluators' judgments automatically without requiring actual human participation in the evaluation process.

Inventive Principle:
Principle #26Copying

2Measurement precision

If human evaluators are used for MOS assessment, then quality can be judged from multiple aspects, but evaluator biases and subjective preferences affect results

Engineering Contradiction:
Improvequality assessment accuracyVSAvoidevaluation consistency
Core Design Contradiction:
Measurement precisionVSReliability

Solution Approach 1:

The patent replaces subjective human judgment with an objective machine learning-based assessment system. The model processes acoustic features through standardized computational algorithms, eliminating variability introduced by different evaluators' preferences, attitudes, and mental states while maintaining consistent and reliable quality assessment across different contexts.

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

3Productivity

If objective methods using acoustic features are used, then evaluation becomes efficient and repeatable, but suitable natural speech references are difficult to obtain

Engineering Contradiction:
Improveevaluation efficiencyVSAvoidevaluation credibility
Core Design Contradiction:
ProductivityVSReliability

Solution Approach 1:

The patent introduces machine learning models as intermediary components that bridge the gap between automated feature extraction and reliable quality assessment. The models process acoustic features and map them to quality scores, enabling efficient automated evaluation while maintaining the credibility and reliability of the assessment through learned relationships from training data.

Inventive Principle:
Principle #24Intermediary (Mediator)

4Ease of manufacture

If bottom-layer acoustic features are used for evaluation, then extraction is straightforward, but high-layer characteristics like naturalness and intonation are difficult to capture

Engineering Contradiction:
Improvefeature extraction simplicityVSAvoidhigh-layer characteristic capture
Core Design Contradiction:
Ease of manufactureVSMeasurement precision

Solution Approach 1:

The patent transitions from analyzing only bottom-layer acoustic features to incorporating high-layer characteristics by processing features through machine learning models. The models operate in a transformed feature space where both acoustic and perceptual characteristics can be simultaneously captured and evaluated, enabling precise measurement of naturalness, intonation, and other high-layer properties.

Inventive Principle:
Principle #17Another dimension (Dimensionality change)

Data Source

PatentUS10950256B2System and method for text-to-speech performance evaluation
Publication Date: 2021.03.16 BAYERISCHE MOTOREN WERKE AG
  • US10950256B2 patent drawing
  • US10950256B2 patent drawing
  • US10950256B2 patent drawing

AI summary

A system and method for text-to-speech performance evaluation are provided. The method (100) for text-to-speech performance evaluation includes providing a plurality of speech samples and scores associated with the respective speech samples (110); extracting acoustic features that influence the associated scores of the respective speech samples from the respective speech samples (120); training a machine learning model by the acoustic features and corresponding scores (130); and evaluating a text-to-speech engine by the trained machine learning model (140).