Text-to-Speech Accent Accuracy via Self-Generated Learning Data
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Conventional text-to-speech systems face challenges in generating natural and accurate accents, requiring extensive expert effort and large amounts of learning data, leading to inconsistent and unnatural synthesized speech due to manual input and limited context consideration.
Innovation Solution
A system comprising a learning data generating unit, frequency data generating unit, and setting unit that recognizes inputted speech, generates learning data associating phrases with readings, computes appearance frequencies, and sets frequency data in a language processing unit to approximate output speech to the inputted speech, enabling efficient generation of high-quality synthesized speech.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Productivity
If statistical information methods are used to determine pronunciation and accent, then processing efficiency is improved, but large amounts of learning data with accurate annotations are required
Solution Approach 1:
The system enables the text-to-speech apparatus to automatically generate and update its own learning data by capturing and storing actual speech outputs and their corresponding text inputs. This self-service mechanism eliminates the need for external experts to manually annotate large amounts of learning data, while still providing sufficient training data for the statistical information method to operate effectively.
Solution Approach 2:
The system performs preliminary data collection and storage by capturing speech outputs and associated text during normal operation, building up a repository of learning data before it is needed for training. This preliminary action ensures that sufficient annotated data is available when training is required, without needing to collect it all at once externally.
2Measurement precision
If experts manually provide accent information for learning data, then accurate pronunciation and accent determination is achieved, but enormous costs and time are required
Solution Approach 1:
The text-to-speech apparatus serves itself by automatically generating learning data with accurate accent information captured from its own speech outputs. The system stores pairs of text inputs and corresponding speech outputs, which inherently contain accurate accent and pronunciation information without requiring external expert annotation.
Solution Approach 2:
The speech output itself acts as an intermediary that carries the accent information. Instead of experts directly annotating text, the system captures the actual spoken output which naturally embodies the correct accents and pronunciations, using this speech data as the intermediary to transfer accurate accent information to the learning data.
3Measurement precision
If manual expert input is used to generate learning data, then accurate accent information is obtained, but inconsistency occurs between manual accents and synthesized speech
Solution Approach 1:
By having the text-to-speech apparatus generate its own learning data from its actual speech outputs, the system ensures inherent consistency between the accent information in the learning data and the synthesized speech. The same system that produces the speech also captures and stores the accent patterns, eliminating the inconsistency that arises when external experts annotate data.
4Measurement precision
If rules are generated by trial-and-error analysis of standard speeches, then appropriate accents can be determined, but various kinds of expert work are required
Solution Approach 1:
The system replaces the mechanical process of manual rule generation by experts with an automated statistical learning process. Instead of experts analyzing standard speeches and creating rules through trial-and-error, the system automatically captures speech data and uses statistical methods to learn accent patterns, substituting manual mechanical work with automated computational processing.
Data Source
AI summary
A system for generating high-quality synthesized text-to-speech includes a learning data generating unit, a frequency data generating unit, and a setting unit. The learning data generating unit recognizes inputted speech, and then generates first learning data in which wordings of phrases are associated with readings thereof. The frequency data generating unit generates, based on the first learning data, frequency data indicating appearance frequencies of both wordings and readings of phrases. The setting unit sets the thus generated frequency data for a language processing unit in order to approximate outputted speech of text-to-speech to the inputted speech. Furthermore, the language processing unit generates, from a wording of text, a reading corresponding to the wording, on the basis of the appearance frequencies.


