Speech Synthesizer Quality Evaluation Model

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing speech synthesizers lack comprehensive and objective methods for evaluating the quality of synthesized speech, making it difficult to determine if the generated speech meets desired purposes.

Innovation Solution

A speech synthesizer using artificial intelligence that includes a database for storing synthesized and correct speech, a processor for comparing speech feature sets, and a speech quality evaluation model to determine quality indices and model parameters, allowing for objective and quantitative evaluation of synthesized speech.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If conventional speech synthesis is used, then speech generation is achieved, but objective and quantitative quality evaluation is insufficient

Engineering Contradiction:
Improvespeech quality evaluationVSAvoidevaluation system
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The patent introduces a speech quality evaluation model as an intermediary component that bridges the gap between synthesized speech and objective quality assessment. This model processes speech features and generates quality indices, enabling quantitative evaluation without requiring complex manual analysis systems.

Inventive Principle:
Principle #24Intermediary (Mediator)

Solution Approach 2:

The patent replaces subjective human listening evaluation with an automated computational evaluation system. The speech quality evaluation model uses machine learning algorithms to objectively assess speech quality based on extracted features, substituting the mechanical process of human perception with an automated digital system.

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

2Reliability

If speech quality evaluation model is introduced, then objective and quantitative evaluation is enabled, but system complexity increases

Engineering Contradiction:
Improvequality assessment accuracyVSAvoidmodel parameters
Core Design Contradiction:
ReliabilityVSDevice complexity

Solution Approach 1:

The patent manages model complexity by dynamically adjusting and optimizing parameters within the speech quality evaluation model. The system trains the model using labeled data to determine optimal parameter values, allowing accurate quality assessment while maintaining manageable complexity through systematic parameter optimization rather than exhaustive model design.

Inventive Principle:
Principle #35Parameter changes

3Measurement precision

If comprehensive quality evaluation is implemented, then speech quality accuracy is improved, but evaluation time increases

Engineering Contradiction:
Improvequality evaluation accuracyVSAvoidevaluation time
Core Design Contradiction:
Measurement precisionVSLoss of time

Solution Approach 1:

The patent performs preliminary feature extraction from speech signals before conducting the actual quality evaluation. By pre-processing the speech data to extract relevant features (such as spectral characteristics, temporal features, and perceptual features), the system prepares the input data in advance, enabling faster and more accurate quality assessment without requiring extensive processing during the evaluation phase.

Inventive Principle:
Principle #10Preliminary action

Data Source

PatentUS11705105B2Speech synthesizer for evaluating quality of synthesized speech using artificial intelligence and method of operating the same
Publication Date: 2023.07.18 LG ELECTRONICS INC
  • US11705105B2 patent drawing
  • US11705105B2 patent drawing
  • US11705105B2 patent drawing

AI summary

A speech synthesizer for evaluating quality of a synthesized speech using artificial intelligence includes a database configured to store a synthesized speech corresponding to text, a correct speech corresponding to the text and a speech quality evaluation model for evaluating the quality of the synthesized speech, and a processor configured to compare a first speech feature set indicating a feature of the synthesized speech and a second speech feature set indicating a feature of the correct speech, acquire a quality evaluation index set including indices used to evaluate the quality of the synthesized speech according to a result of comparison, and determine weights as model parameters of the speech quality evaluation model using the acquired quality evaluation index set and the speech quality evaluation model.