Speech Synthesis Feedback Loop for Accurate Text-to-Speech Data Generation
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Speech recognition models struggle to learn accurately from simulation data generated without pronunciation information, as inaccurate pronunciation estimation leads to mismatched speech and text pairs, hindering the learning process.
Innovation Solution
A data generation apparatus that includes a speech synthesis unit, a speech recognition unit, and a matching processing unit, which estimates pronunciation and accent from text, generates speech data, performs speech recognition, and iteratively refines pronunciation candidates to achieve high matching degrees between original and recognition texts, thereby creating reliable simulation data.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Productivity
If speech synthesis is performed from text without pronunciation information, then simulation data can be generated, but pronunciation accuracy deteriorates leading to inaccurate synthetic speech
Solution Approach 1:
The system performs speech recognition on the synthesized speech and compares the recognized text with the original text. This feedback loop identifies pronunciation errors by detecting mismatches, allowing the system to iteratively improve pronunciation accuracy until the recognized text matches the original text.
Solution Approach 2:
The system performs speech synthesis in advance and then uses speech recognition to evaluate the pronunciation quality before the data is used for training. This preliminary evaluation allows identification and correction of pronunciation errors before the simulation data is finalized for model training.
2Device complexity
If pronunciation information is not provided in the original text, then data processing is simplified, but speech recognition model learning accuracy deteriorates
Solution Approach 1:
The system automatically estimates pronunciation information through speech synthesis and recognition without requiring manual annotation. The speech recognition model itself serves to generate the pronunciation data needed for training, eliminating the need for external pronunciation information while maintaining high learning accuracy.
3Speed
If speech synthesis is performed without accurate pronunciation estimation, then processing speed is improved, but synthetic speech quality deteriorates
Solution Approach 1:
The system uses speech recognition as a feedback mechanism to evaluate synthetic speech quality. By comparing the recognized text with the original text, the system identifies pronunciation errors and iteratively improves the synthesis process, ensuring high speech quality without excessive manual intervention.
Data Source
AI summary
According to one embodiment, the data generation apparatus includes a speech synthesis unit, a speech recognition unit, a matching processing unit, and a dataset generation unit. The speech synthesis unit generates speech data from an original text. The speech recognition unit generates a recognition text by speech recognition from the speech data. The matching processing unit performs matching between the original text and the recognition text. The dataset generation unit generates a dataset in such a manner where the speech data, from which the recognition text satisfying a certain condition for a matching degree relative to the original text is generated, is associated with the original text, based on a matching result.


