Text-to-Speech Evaluation Using Machine Learning Acoustic Features
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current text-to-speech (TTS) performance evaluation methods are time-consuming, labor-intensive, and subjective, relying on human evaluators for Mean Opinion Score (MOS) assessments, which lack generalization and are affected by evaluator biases, while objective methods struggle with suitable natural speech references and extracting high-layer characteristics.
Innovation Solution
A system and method using machine learning models, such as Support Vector Machines (SVM) and Deep Neural Networks (DNN), to evaluate TTS systems by extracting relevant acoustic features from speech samples and training models with corresponding scores, applying sub-space decomposition methods like Linear Discriminant Analysis (LDA) to select influential features.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If human evaluators are used for MOS assessment, then evaluation can capture subjective quality aspects, but the process becomes time-consuming and labor-intensive
Solution Approach 1:
The patent replaces the mechanical human evaluation system with an automated machine learning-based evaluation system. The machine learning model processes acoustic features extracted from speech samples to automatically predict quality scores, eliminating the need for manual listener assessments while maintaining evaluation accuracy.
Solution Approach 2:
The patent creates a computational model that copies and replicates the evaluation capabilities of human listeners. The machine learning model learns from training data consisting of acoustic features and corresponding quality scores, enabling it to replicate human evaluators' judgments automatically without requiring actual human participation in the evaluation process.
2Measurement precision
If human evaluators are used for MOS assessment, then quality can be judged from multiple aspects, but evaluator biases and subjective preferences affect results
Solution Approach 1:
The patent replaces subjective human judgment with an objective machine learning-based assessment system. The model processes acoustic features through standardized computational algorithms, eliminating variability introduced by different evaluators' preferences, attitudes, and mental states while maintaining consistent and reliable quality assessment across different contexts.
3Productivity
If objective methods using acoustic features are used, then evaluation becomes efficient and repeatable, but suitable natural speech references are difficult to obtain
Solution Approach 1:
The patent introduces machine learning models as intermediary components that bridge the gap between automated feature extraction and reliable quality assessment. The models process acoustic features and map them to quality scores, enabling efficient automated evaluation while maintaining the credibility and reliability of the assessment through learned relationships from training data.
4Ease of manufacture
If bottom-layer acoustic features are used for evaluation, then extraction is straightforward, but high-layer characteristics like naturalness and intonation are difficult to capture
Solution Approach 1:
The patent transitions from analyzing only bottom-layer acoustic features to incorporating high-layer characteristics by processing features through machine learning models. The models operate in a transformed feature space where both acoustic and perceptual characteristics can be simultaneously captured and evaluated, enabling precise measurement of naturalness, intonation, and other high-layer properties.
Data Source
AI summary
A system and method for text-to-speech performance evaluation are provided. The method (100) for text-to-speech performance evaluation includes providing a plurality of speech samples and scores associated with the respective speech samples (110); extracting acoustic features that influence the associated scores of the respective speech samples from the respective speech samples (120); training a machine learning model by the acoustic features and corresponding scores (130); and evaluating a text-to-speech engine by the trained machine learning model (140).


