Speech Synthesis Model with Automated Quality Evaluation
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing speech synthesis technologies rely on manual evaluation, which is inefficient and prone to errors, especially in scenarios requiring automation, and often result in incorrect text processing leading to speech synthesis errors.
Innovation Solution
A method and system for automatically synthesizing speech using a speech synthesis model comprising an embedding layer, a speech synthesis layer, and a position layer, with the model being trained when an evaluation index meets a preset condition, ensuring improved synthesis efficiency and accuracy.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Productivity
If manual evaluation is used for speech synthesis, then evaluation accuracy can be maintained, but synthesis efficiency deteriorates and automation requirements cannot be met
Solution Approach 1:
The system performs self-evaluation by automatically comparing synthesized speech with reference speech using objective evaluation indexes (MOS, PESQ, STOI), eliminating the need for manual human evaluation and enabling automated quality assessment
Solution Approach 2:
The system implements a feedback mechanism where evaluation results are used to automatically trigger re-synthesis operations, creating a closed-loop system that continuously improves synthesis quality through automated iteration
2Reliability
If text is processed with speech synthesis model, then speech can be generated, but text processing errors occur leading to speech synthesis errors
Solution Approach 1:
The system performs preliminary evaluation of the synthesized speech against reference speech before final output, detecting potential errors in text processing early in the synthesis pipeline
Solution Approach 2:
The evaluation mechanism provides feedback on text processing quality by comparing synthesized speech characteristics with reference speech, enabling detection and correction of text processing errors
3Manufacturing precision
If speech synthesis model is continuously trained, then synthesis quality improves, but system complexity and maintenance costs increase
Solution Approach 1:
The system implements dynamic training triggers based on evaluation results, automatically initiating re-training only when synthesis quality deteriorates below threshold levels, rather than continuous training
Solution Approach 2:
The system uses evaluation feedback to dynamically control the training process, adjusting when and how training occurs based on actual synthesis performance rather than following a fixed schedule
Data Source
AI summary
The present disclosure discloses a method for synthesizing a speech. The method includes generating the speech based on a text with a speech synthesis model, wherein the speech synthesis model includes an embedding layer, a speech synthesis layer, and a position layer; and training the speech synthesis model when an evaluation index meets a preset condition, wherein the evaluation index includes one or more quality indexes determined based on at least a part of the text and at least a part of the speech.


