Hybrid Text-to-Speech Voice Data Compression

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Text-to-speech (TTS) systems face challenges in reducing storage requirements for voice data, particularly on devices with limited capacity, as uncompressed voice data can exceed available storage, necessitating efficient compression methods without compromising audio quality.

Innovation Solution

The use of time domain compression followed by perceptual compression techniques to reduce storage space, with varying compression ratios applied based on linguistic and acoustic features of speech segments, allowing for customizable decompression for accurate playback.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Reliability

If voice data is stored uncompressed, then audio quality is maintained, but storage requirements exceed available device capacity

Engineering Contradiction:
Improveaudio qualityVSAvoidstorage requirements
Core Design Contradiction:
ReliabilityVSQuantity of substance

Solution Approach 1:

The patent segments voice data into distinct phoneme categories (voiced, unvoiced, and pause segments) and applies different compression strategies to each type. This segmentation allows the system to preserve audio quality for critical segments while aggressively compressing less critical segments, resolving the contradiction between maintaining overall audio quality and reducing total storage requirements.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent applies local quality by using variable compression ratios tailored to specific phoneme types. Voiced segments receive higher quality preservation with lower compression ratios, while unvoiced and pause segments undergo more aggressive compression. This localized quality approach ensures that perceptually important elements maintain fidelity while reducing overall storage burden.

Inventive Principle:
Principle #3Local quality

2Quantity of substance

If aggressive compression is applied to reduce storage space, then storage requirements are reduced, but audio quality deteriorates

Engineering Contradiction:
Improvestorage spaceVSAvoidaudio quality
Core Design Contradiction:
Quantity of substanceVSReliability

Solution Approach 1:

By dividing the audio stream into distinct phoneme segments and categorizing them by type (voiced, unvoiced, pause), the system can apply differentiated compression strategies. This prevents uniform aggressive compression from degrading overall audio quality, as critical voiced segments receive gentler treatment while less critical segments are compressed more heavily.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent changes compression parameters based on phoneme type characteristics. Different compression ratios, encoding schemes, and quality thresholds are applied to voiced versus unvoiced segments. This parameter adaptation allows the system to achieve higher overall compression ratios while maintaining perceptual quality by adjusting parameters to match the inherent characteristics of each segment type.

Inventive Principle:
Principle #35Parameter changes

3Device complexity

If uniform compression is applied to all speech segments, then processing is simplified, but storage efficiency is reduced

Engineering Contradiction:
Improveprocessing complexityVSAvoidstorage efficiency
Core Design Contradiction:
Device complexityVSQuantity of substance

Solution Approach 1:

The patent introduces segmentation based on phoneme classification, which adds processing structure rather than complexity. By automatically categorizing segments into voiced, unvoiced, and pause types, the system creates a manageable framework that enables efficient storage optimization without requiring complex manual intervention or overly sophisticated processing algorithms.

Inventive Principle:
Principle #1Segmentation

Data Source

PatentUS9064489B2Hybrid compression of text-to-speech voice data
Publication Date: 2015.06.23 AMAZON TECH INC
  • US9064489B2 patent drawing
  • US9064489B2 patent drawing
  • US9064489B2 patent drawing

AI summary

Recorded or synthesized speech segments of text-to-speech (TTS) systems may be compressed though the use of both time domain compression and perceptual compression techniques. The twice-compressed recording may be separated into speech segments corresponding to words or subword units for use in a TTS system. The compression rate of time domain compression, and the ratio of time domain compression to perceptual compression, may be modified for any speech segment. The compression amount or ratio may be determined based on linguistic or acoustic features of the word or subword unit that the speech segment represents. Differing compression amounts and ratios may be applied to portions of a single speech segment.