Speech Synthesis Feedback Loop for Accurate Text-to-Speech Data Generation

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Speech recognition models struggle to learn accurately from simulation data generated without pronunciation information, as inaccurate pronunciation estimation leads to mismatched speech and text pairs, hindering the learning process.

Innovation Solution

A data generation apparatus that includes a speech synthesis unit, a speech recognition unit, and a matching processing unit, which estimates pronunciation and accent from text, generates speech data, performs speech recognition, and iteratively refines pronunciation candidates to achieve high matching degrees between original and recognition texts, thereby creating reliable simulation data.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Productivity

If speech synthesis is performed from text without pronunciation information, then simulation data can be generated, but pronunciation accuracy deteriorates leading to inaccurate synthetic speech

Engineering Contradiction:
Improvesimulation data generation capabilityVSAvoidpronunciation accuracy
Core Design Contradiction:
ProductivityVSManufacturing precision

Solution Approach 1:

The system performs speech recognition on the synthesized speech and compares the recognized text with the original text. This feedback loop identifies pronunciation errors by detecting mismatches, allowing the system to iteratively improve pronunciation accuracy until the recognized text matches the original text.

Inventive Principle:
Principle #23Feedback

Solution Approach 2:

The system performs speech synthesis in advance and then uses speech recognition to evaluate the pronunciation quality before the data is used for training. This preliminary evaluation allows identification and correction of pronunciation errors before the simulation data is finalized for model training.

Inventive Principle:
Principle #10Preliminary action

2Device complexity

If pronunciation information is not provided in the original text, then data processing is simplified, but speech recognition model learning accuracy deteriorates

Engineering Contradiction:
Improvedata processing complexityVSAvoidmodel learning accuracy
Core Design Contradiction:
Device complexityVSReliability

Solution Approach 1:

The system automatically estimates pronunciation information through speech synthesis and recognition without requiring manual annotation. The speech recognition model itself serves to generate the pronunciation data needed for training, eliminating the need for external pronunciation information while maintaining high learning accuracy.

Inventive Principle:
Principle #25Self-service

3Speed

If speech synthesis is performed without accurate pronunciation estimation, then processing speed is improved, but synthetic speech quality deteriorates

Engineering Contradiction:
Improveprocessing speedVSAvoidsynthetic speech quality
Core Design Contradiction:
SpeedVSManufacturing precision

Solution Approach 1:

The system uses speech recognition as a feedback mechanism to evaluate synthetic speech quality. By comparing the recognized text with the original text, the system identifies pronunciation errors and iteratively improves the synthesis process, ensuring high speech quality without excessive manual intervention.

Inventive Principle:
Principle #23Feedback

Data Source

PatentUS11694028B2Data generation apparatus and data generation method that generate recognition text from speech data
Publication Date: 2023.07.04 KK TOSHIBA
  • US11694028B2 patent drawing
  • US11694028B2 patent drawing
  • US11694028B2 patent drawing

AI summary

According to one embodiment, the data generation apparatus includes a speech synthesis unit, a speech recognition unit, a matching processing unit, and a dataset generation unit. The speech synthesis unit generates speech data from an original text. The speech recognition unit generates a recognition text by speech recognition from the speech data. The matching processing unit performs matching between the original text and the recognition text. The dataset generation unit generates a dataset in such a manner where the speech data, from which the recognition text satisfying a certain condition for a matching degree relative to the original text is generated, is associated with the original text, based on a matching result.