Distributed Speech Unit Inventory for TTS Latency

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Distributed text-to-speech (TTS) systems face high network load and latency due to reliance on centralized servers for processing, which can result in delays and prevent TTS processing if a network connection is unavailable.

Innovation Solution

Implementing a local device with a smaller subset of the centralized unit database to perform localized TTS processing, which checks for available units locally and communicates with a remote server only when necessary to offload processing and reduce bandwidth usage.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If a centralized server is used for TTS processing, then high-quality speech synthesis is achieved, but network load and latency increase

Engineering Contradiction:
Improvespeech synthesis qualityVSAvoidprocessing latency
Core Design Contradiction:
Measurement precisionVSLoss of time

Solution Approach 1:

The patent segments the centralized TTS system into distributed components across multiple devices. Each device maintains a local speech unit inventory (subset of the full corpus), allowing local TTS processing without requiring constant server communication. This segmentation reduces network latency while maintaining speech quality through distributed processing.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent implements local quality by enabling each device to perform TTS processing using its own local speech unit inventory. This allows speech synthesis to be performed locally with high quality for frequently encountered text, eliminating the need for remote server processing and reducing latency.

Inventive Principle:
Principle #3Local quality

2Adaptability or versatility

If a centralized server is used for TTS processing, then comprehensive speech units are available, but server load increases

Engineering Contradiction:
Improvespeech unit availabilityVSAvoidserver processing capacity
Core Design Contradiction:
Adaptability or versatilityVSProductivity

Solution Approach 1:

The patent divides the comprehensive speech unit corpus into distributed local inventories across multiple devices. Each device stores a subset of speech units locally, reducing the processing burden on any single server while maintaining comprehensive coverage across the distributed system.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent implements partial action by having each device maintain only a subset of the complete speech unit corpus locally. This partial local inventory is sufficient for most TTS operations, reducing server load while maintaining versatility for frequently encountered text.

Inventive Principle:
Principle #16Partial or excessive action

3Loss of energy

If a local device performs TTS processing with a smaller unit database, then bandwidth usage decreases, but speech synthesis quality may be reduced

Engineering Contradiction:
Improvebandwidth usageVSAvoidspeech synthesis quality
Core Design Contradiction:
Loss of energyVSMeasurement precision

Solution Approach 1:

The patent applies preliminary action by pre-populating each device's local speech unit inventory with frequently used speech units before TTS processing is needed. This preliminary local storage enables high-quality speech synthesis for common text without requiring bandwidth-intensive remote retrieval during operation.

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The patent optimizes local quality by configuring each device's local speech unit inventory to contain the most frequently encountered speech units. This localized optimization ensures high speech synthesis quality for typical usage scenarios while minimizing bandwidth requirements.

Inventive Principle:
Principle #3Local quality

Data Source

PatentEP2943950B8Distributed speech unit inventory for TTS systems
Publication Date: 2017.03.08 AMAZON TECH INC
  • EP2943950B8 patent drawing

AI summary

In a text-to-speech (TTS) system, a database including sample speech units for unit selection may be configured for use by a local device. The local unit database may be created from a more comprehensive unit database. The local unit database may include units which provide sufficient TTS results for frequently input text. Speech synthesis may then be performed by concatenating locally available units with units from a remote device including the comprehensive unit database. Aspects of the speech synthesis may be performed by the remote device and/or the local device.