Text-to-Speech Synthesis Using Pronunciation Databases

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Publishers face high costs and impracticality in converting printed materials into audio formats, especially for dynamic content like newspapers that require frequent updates, due to the need for professional recording and resource-intensive processes.

Innovation Solution

Implementing a text-to-speech algorithm that uses pronunciation databases and voice data to automatically synthesize digital text into audio, allowing publishers to provide dynamic audio versions of their content seamlessly and conveniently, with options for multiple voices and customizable pronunciation.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Reliability

If professional recording is used to convert printed materials into audio formats, then audio quality and professionalism are improved, but costs and resource commitments increase significantly

Engineering Contradiction:
Improveaudio qualityVSAvoidcost and resource commitment
Core Design Contradiction:
ReliabilityVSEase of manufacture

Solution Approach 1:

The system enables publishers to self-generate audio content using automated text-to-speech algorithms without requiring professional recording services. The electronic device synthesizes speech from digital text using pronunciation databases and voice data, allowing the publisher to produce audio versions independently, thereby eliminating the need for external professional resources while maintaining acceptable audio quality

Inventive Principle:
Principle #25Self-service

Solution Approach 2:

The patent replaces the mechanical process of professional human recording with an automated computational system. The text-to-speech algorithm substitutes human performers and recording equipment with software-based synthesis using pronunciation databases and voice data, dramatically reducing resource requirements while producing audio output

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

2Manufacturing precision

If professional recording is used for dynamic content like newspapers, then content accuracy is improved, but the impracticality and resource requirements worsen due to frequent updates

Engineering Contradiction:
Improvecontent accuracyVSAvoidupdate efficiency
Core Design Contradiction:
Manufacturing precisionVSProductivity

Solution Approach 1:

The system is designed to dynamically generate audio content from digital text sources. When content updates occur, the text-to-speech algorithm automatically re-synthesizes the audio from the updated text, allowing frequent updates without requiring new recording sessions. This dynamic approach maintains content accuracy while enabling rapid production cycles suitable for newspapers and periodicals

Inventive Principle:
Principle #15Dynamics

Solution Approach 2:

The system performs preliminary synthesis of audio from digital text, storing pronunciation data and voice characteristics in advance. When updates are needed, only the text changes require re-synthesis while leveraging pre-loaded pronunciation databases, making frequent updates practical and efficient

Inventive Principle:
Principle #10Preliminary action

3Productivity

If text-to-speech algorithm is used to synthesize digital text into audio, then costs are reduced and productivity is improved, but device complexity increases due to pronunciation databases and voice data requirements

Engineering Contradiction:
Improvecontent production efficiencyVSAvoidsystem complexity
Core Design Contradiction:
ProductivityVSDevice complexity

Solution Approach 1:

The system segments the text-to-speech functionality into separate modular components: pronunciation databases, voice data files, and synthesis algorithms. These can be independently managed, updated, and stored, reducing the complexity burden on any single component while maintaining high productivity through their coordinated operation

Inventive Principle:
Principle #1Segmentation

Data Source

PatentUS8990087B1Providing text to speech from digital content on an electronic device
Publication Date: 2015.03.24 AMAZON TECH INC
  • US8990087B1 patent drawing
  • US8990087B1 patent drawing
  • US8990087B1 patent drawing

AI summary

A method for providing text to speech from digital content in an electronic device is described. Digital content including a plurality of words and a pronunciation database is received. Pronunciation instructions are determined for the word using the digital content. Audio or speech is played for the word using the pronunciation instructions. As a result, the method provides text to speech on the electronic device based on the digital content.