Text-to-Speech Evaluation Model Using Machine Learning

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current text-to-speech (TTS) synthesis evaluation methods rely on subjective human assessment, which is time-consuming, labor-intensive, and lacks automation, making it difficult to efficiently and objectively select the best TTS products or assess performance improvements.

Innovation Solution

A supervised machine learning approach is employed to automatically evaluate TTS engines through data sampling and rating, followed by speech modeling and evaluation, using a combination of human and TTS speech samples, and objective scoring methods like MOS, DAM, and CT, to create a reusable speech model for evaluating TTS performance.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Reliability

If subject evaluation methods (MOS, DAM, CT) are used to evaluate TTS performance, then evaluation accuracy and reliability are improved, but time consumption and labor intensity increase significantly

Engineering Contradiction:
Improveevaluation reliabilityVSAvoidtime consumption
Core Design Contradiction:
ReliabilityVSLoss of time

Solution Approach 1:

The patent creates a speech model that copies the evaluation capabilities of human listeners through machine learning. The model is trained on speech samples with associated evaluation scores, enabling it to replicate human subjective evaluation processes automatically without requiring actual human participants for each evaluation task.

Inventive Principle:
Principle #26Copying

Solution Approach 2:

The patent replaces the mechanical system of human subjective evaluation with an automated computational system. The speech model uses objective acoustic feature extraction and machine learning algorithms to substitute the manual, time-consuming process of human listening and rating, while maintaining evaluation reliability.

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

2Measurement precision

If subject evaluation methods are used to evaluate TTS performance, then evaluation accuracy is improved, but labor intensity and cost increase

Engineering Contradiction:
Improveevaluation accuracyVSAvoidlabor intensity
Core Design Contradiction:
Measurement precisionVSEase of manufacture

Solution Approach 1:

The speech evaluation system is self-service in that once trained, it can autonomously evaluate TTS speech without requiring human intervention for each evaluation. The model automatically extracts features, compares speech samples, and generates evaluation scores, eliminating the need for continuous human labor in the evaluation process.

Inventive Principle:
Principle #25Self-service

Solution Approach 2:

The system copies human evaluation expertise into a reusable computational model. By training on labeled speech data with human evaluation scores, the model captures and replicates human evaluation patterns, enabling accurate automated assessment without ongoing human involvement.

Inventive Principle:
Principle #26Copying

3Reliability

If subjective tests with multiple listeners are conducted to reduce result uncertainty, then evaluation reliability is improved, but time consumption and cost increase

Engineering Contradiction:
Improveresult repeatabilityVSAvoidevaluation efficiency
Core Design Contradiction:
ReliabilityVSProductivity

Solution Approach 1:

The speech model copies the collective evaluation judgment of multiple listeners into a single computational entity. By training on data from multiple evaluated speech samples, the model internalizes diverse listener perspectives and can reproduce consistent, repeatable results without requiring multiple actual listeners for each evaluation.

Inventive Principle:
Principle #26Copying

Solution Approach 2:

The evaluation system performs preliminary action by pre-training the speech model on comprehensive speech data with evaluation scores before actual evaluation tasks. This preliminary training phase captures the variability and consensus of multiple listeners, enabling the model to deliver repeatable results efficiently during deployment without requiring repeated human testing.

Inventive Principle:
Principle #10Preliminary action

Data Source

PatentEP3061086B1Text-to-speech performance evaluation
Publication Date: 2019.10.23 BAYERISCHE MOTOREN WERKE AG
  • EP3061086B1 patent drawingFigure 1~2
  • EP3061086B1 patent drawingFigure 3
  • EP3061086B1 patent drawingFigure 4

AI summary

The present invention provides systems and methods for text-to-speech performance evaluation. In an exemplary embodiment, A method for text-to-speech performance evaluation, comprising: providing a plurality of speech samples and scores associated with the respective speech samples; establishing a speech model based on the plurality of speech samples and the corresponding scores; and evaluating a TTS engine by the speech model. The present invention only requires one person to generate a standard speech model at the beginning stage, and this speech model can be repetitively used for test and evaluation of different TTS synthesis engines. The proposed solution in this invention largely decreases the required time and labour cost.