Speech Synthesizer Quality Evaluation Model
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing speech synthesizers lack comprehensive and objective methods for evaluating the quality of synthesized speech, making it difficult to determine if the generated speech meets desired purposes.
Innovation Solution
A speech synthesizer using artificial intelligence that includes a database for storing synthesized and correct speech, a processor for comparing speech feature sets, and a speech quality evaluation model to determine quality indices and model parameters, allowing for objective and quantitative evaluation of synthesized speech.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If conventional speech synthesis is used, then speech generation is achieved, but objective and quantitative quality evaluation is insufficient
Solution Approach 1:
The patent introduces a speech quality evaluation model as an intermediary component that bridges the gap between synthesized speech and objective quality assessment. This model processes speech features and generates quality indices, enabling quantitative evaluation without requiring complex manual analysis systems.
Solution Approach 2:
The patent replaces subjective human listening evaluation with an automated computational evaluation system. The speech quality evaluation model uses machine learning algorithms to objectively assess speech quality based on extracted features, substituting the mechanical process of human perception with an automated digital system.
2Reliability
If speech quality evaluation model is introduced, then objective and quantitative evaluation is enabled, but system complexity increases
Solution Approach 1:
The patent manages model complexity by dynamically adjusting and optimizing parameters within the speech quality evaluation model. The system trains the model using labeled data to determine optimal parameter values, allowing accurate quality assessment while maintaining manageable complexity through systematic parameter optimization rather than exhaustive model design.
3Measurement precision
If comprehensive quality evaluation is implemented, then speech quality accuracy is improved, but evaluation time increases
Solution Approach 1:
The patent performs preliminary feature extraction from speech signals before conducting the actual quality evaluation. By pre-processing the speech data to extract relevant features (such as spectral characteristics, temporal features, and perceptual features), the system prepares the input data in advance, enabling faster and more accurate quality assessment without requiring extensive processing during the evaluation phase.
Data Source
AI summary
A speech synthesizer for evaluating quality of a synthesized speech using artificial intelligence includes a database configured to store a synthesized speech corresponding to text, a correct speech corresponding to the text and a speech quality evaluation model for evaluating the quality of the synthesized speech, and a processor configured to compare a first speech feature set indicating a feature of the synthesized speech and a second speech feature set indicating a feature of the correct speech, acquire a quality evaluation index set including indices used to evaluate the quality of the synthesized speech according to a result of comparison, and determine weights as model parameters of the speech quality evaluation model using the acquired quality evaluation index set and the speech quality evaluation model.


