Hybrid Text-to-Speech Voice Data Compression
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Text-to-speech (TTS) systems face challenges in reducing storage requirements for voice data, particularly on devices with limited capacity, as uncompressed voice data can exceed available storage, necessitating efficient compression methods without compromising audio quality.
Innovation Solution
The use of time domain compression followed by perceptual compression techniques to reduce storage space, with varying compression ratios applied based on linguistic and acoustic features of speech segments, allowing for customizable decompression for accurate playback.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Reliability
If voice data is stored uncompressed, then audio quality is maintained, but storage requirements exceed available device capacity
Solution Approach 1:
The patent segments voice data into distinct phoneme categories (voiced, unvoiced, and pause segments) and applies different compression strategies to each type. This segmentation allows the system to preserve audio quality for critical segments while aggressively compressing less critical segments, resolving the contradiction between maintaining overall audio quality and reducing total storage requirements.
Solution Approach 2:
The patent applies local quality by using variable compression ratios tailored to specific phoneme types. Voiced segments receive higher quality preservation with lower compression ratios, while unvoiced and pause segments undergo more aggressive compression. This localized quality approach ensures that perceptually important elements maintain fidelity while reducing overall storage burden.
2Quantity of substance
If aggressive compression is applied to reduce storage space, then storage requirements are reduced, but audio quality deteriorates
Solution Approach 1:
By dividing the audio stream into distinct phoneme segments and categorizing them by type (voiced, unvoiced, pause), the system can apply differentiated compression strategies. This prevents uniform aggressive compression from degrading overall audio quality, as critical voiced segments receive gentler treatment while less critical segments are compressed more heavily.
Solution Approach 2:
The patent changes compression parameters based on phoneme type characteristics. Different compression ratios, encoding schemes, and quality thresholds are applied to voiced versus unvoiced segments. This parameter adaptation allows the system to achieve higher overall compression ratios while maintaining perceptual quality by adjusting parameters to match the inherent characteristics of each segment type.
3Device complexity
If uniform compression is applied to all speech segments, then processing is simplified, but storage efficiency is reduced
Solution Approach 1:
The patent introduces segmentation based on phoneme classification, which adds processing structure rather than complexity. By automatically categorizing segments into voiced, unvoiced, and pause types, the system creates a manageable framework that enables efficient storage optimization without requiring complex manual intervention or overly sophisticated processing algorithms.
Data Source
AI summary
Recorded or synthesized speech segments of text-to-speech (TTS) systems may be compressed though the use of both time domain compression and perceptual compression techniques. The twice-compressed recording may be separated into speech segments corresponding to words or subword units for use in a TTS system. The compression rate of time domain compression, and the ratio of time domain compression to perceptual compression, may be modified for any speech segment. The compression amount or ratio may be determined based on linguistic or acoustic features of the word or subword unit that the speech segment represents. Differing compression amounts and ratios may be applied to portions of a single speech segment.


