Distributed Speech Unit Inventory for TTS Latency
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Distributed text-to-speech (TTS) systems face high network load and latency due to reliance on centralized servers for processing, which can result in delays and prevent TTS processing if a network connection is unavailable.
Innovation Solution
Implementing a local device with a smaller subset of the centralized unit database to perform localized TTS processing, which checks for available units locally and communicates with a remote server only when necessary to offload processing and reduce bandwidth usage.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If a centralized server is used for TTS processing, then high-quality speech synthesis is achieved, but network load and latency increase
Solution Approach 1:
The patent segments the centralized TTS system into distributed components across multiple devices. Each device maintains a local speech unit inventory (subset of the full corpus), allowing local TTS processing without requiring constant server communication. This segmentation reduces network latency while maintaining speech quality through distributed processing.
Solution Approach 2:
The patent implements local quality by enabling each device to perform TTS processing using its own local speech unit inventory. This allows speech synthesis to be performed locally with high quality for frequently encountered text, eliminating the need for remote server processing and reducing latency.
2Adaptability or versatility
If a centralized server is used for TTS processing, then comprehensive speech units are available, but server load increases
Solution Approach 1:
The patent divides the comprehensive speech unit corpus into distributed local inventories across multiple devices. Each device stores a subset of speech units locally, reducing the processing burden on any single server while maintaining comprehensive coverage across the distributed system.
Solution Approach 2:
The patent implements partial action by having each device maintain only a subset of the complete speech unit corpus locally. This partial local inventory is sufficient for most TTS operations, reducing server load while maintaining versatility for frequently encountered text.
3Loss of energy
If a local device performs TTS processing with a smaller unit database, then bandwidth usage decreases, but speech synthesis quality may be reduced
Solution Approach 1:
The patent applies preliminary action by pre-populating each device's local speech unit inventory with frequently used speech units before TTS processing is needed. This preliminary local storage enables high-quality speech synthesis for common text without requiring bandwidth-intensive remote retrieval during operation.
Solution Approach 2:
The patent optimizes local quality by configuring each device's local speech unit inventory to contain the most frequently encountered speech units. This localized optimization ensures high speech synthesis quality for typical usage scenarios while minimizing bandwidth requirements.
Data Source
AI summary
In a text-to-speech (TTS) system, a database including sample speech units for unit selection may be configured for use by a local device. The local unit database may be created from a more comprehensive unit database. The local unit database may include units which provide sufficient TTS results for frequently input text. Speech synthesis may then be performed by concatenating locally available units with units from a remote device including the comprehensive unit database. Aspects of the speech synthesis may be performed by the remote device and/or the local device.
