Text-to-Speech Accent Accuracy via Self-Generated Learning Data

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Conventional text-to-speech systems face challenges in generating natural and accurate accents, requiring extensive expert effort and large amounts of learning data, leading to inconsistent and unnatural synthesized speech due to manual input and limited context consideration.

Innovation Solution

A system comprising a learning data generating unit, frequency data generating unit, and setting unit that recognizes inputted speech, generates learning data associating phrases with readings, computes appearance frequencies, and sets frequency data in a language processing unit to approximate output speech to the inputted speech, enabling efficient generation of high-quality synthesized speech.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Productivity

If statistical information methods are used to determine pronunciation and accent, then processing efficiency is improved, but large amounts of learning data with accurate annotations are required

Engineering Contradiction:
Improveprocessing efficiencyVSAvoidamount of learning data
Core Design Contradiction:
ProductivityVSQuantity of substance

Solution Approach 1:

The system enables the text-to-speech apparatus to automatically generate and update its own learning data by capturing and storing actual speech outputs and their corresponding text inputs. This self-service mechanism eliminates the need for external experts to manually annotate large amounts of learning data, while still providing sufficient training data for the statistical information method to operate effectively.

Inventive Principle:
Principle #25Self-service

Solution Approach 2:

The system performs preliminary data collection and storage by capturing speech outputs and associated text during normal operation, building up a repository of learning data before it is needed for training. This preliminary action ensures that sufficient annotated data is available when training is required, without needing to collect it all at once externally.

Inventive Principle:
Principle #10Preliminary action

2Measurement precision

If experts manually provide accent information for learning data, then accurate pronunciation and accent determination is achieved, but enormous costs and time are required

Engineering Contradiction:
Improveaccuracy of accent determinationVSAvoidtime for expert annotation
Core Design Contradiction:
Measurement precisionVSLoss of time

Solution Approach 1:

The text-to-speech apparatus serves itself by automatically generating learning data with accurate accent information captured from its own speech outputs. The system stores pairs of text inputs and corresponding speech outputs, which inherently contain accurate accent and pronunciation information without requiring external expert annotation.

Inventive Principle:
Principle #25Self-service

Solution Approach 2:

The speech output itself acts as an intermediary that carries the accent information. Instead of experts directly annotating text, the system captures the actual spoken output which naturally embodies the correct accents and pronunciations, using this speech data as the intermediary to transfer accurate accent information to the learning data.

Inventive Principle:
Principle #24Intermediary (Mediator)

3Measurement precision

If manual expert input is used to generate learning data, then accurate accent information is obtained, but inconsistency occurs between manual accents and synthesized speech

Engineering Contradiction:
Improveaccuracy of accent informationVSAvoidconsistency between accent data and synthesized speech
Core Design Contradiction:
Measurement precisionVSStability of the object's composition

Solution Approach 1:

By having the text-to-speech apparatus generate its own learning data from its actual speech outputs, the system ensures inherent consistency between the accent information in the learning data and the synthesized speech. The same system that produces the speech also captures and stores the accent patterns, eliminating the inconsistency that arises when external experts annotate data.

Inventive Principle:
Principle #25Self-service

4Measurement precision

If rules are generated by trial-and-error analysis of standard speeches, then appropriate accents can be determined, but various kinds of expert work are required

Engineering Contradiction:
Improveaccuracy of accent generationVSAvoidcomplexity of rule generation process
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The system replaces the mechanical process of manual rule generation by experts with an automated statistical learning process. Instead of experts analyzing standard speeches and creating rules through trial-and-error, the system automatically captures speech data and uses statistical methods to learn accent patterns, substituting manual mechanical work with automated computational processing.

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

Data Source

PatentUS7921014B2System and method for supporting text-to-speech
Publication Date: 2011.04.05 CERENCE OPERATING CO
  • US7921014B2 patent drawing
  • US7921014B2 patent drawing
  • US7921014B2 patent drawing

AI summary

A system for generating high-quality synthesized text-to-speech includes a learning data generating unit, a frequency data generating unit, and a setting unit. The learning data generating unit recognizes inputted speech, and then generates first learning data in which wordings of phrases are associated with readings thereof. The frequency data generating unit generates, based on the first learning data, frequency data indicating appearance frequencies of both wordings and readings of phrases. The setting unit sets the thus generated frequency data for a language processing unit in order to approximate outputted speech of text-to-speech to the inputted speech. Furthermore, the language processing unit generates, from a wording of text, a reading corresponding to the wording, on the basis of the appearance frequencies.