Speech Synthesis Phase Spectral Modeling

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing text-to-speech (TTS) systems face challenges in generating high-quality speech with limited memory resources, as they require large databases to avoid auditory discontinuities, which is impractical for variable text inputs or devices with limited memory.

Innovation Solution

The use of spectral modeling that incorporates both amplitude and phase information for speech segments, allowing for reduced database size while maintaining high-quality output, by distinguishing between voiced and unvoiced frames and applying different analysis techniques to model parameters, including phase information for clicks to enhance sound quality.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Reliability

If a large speech database is used to avoid auditory discontinuities and produce high-quality speech, then speech quality is improved, but memory resource requirements increase

Engineering Contradiction:
Improvespeech qualityVSAvoiddatabase size
Core Design Contradiction:
ReliabilityVSQuantity of substance

Solution Approach 1:

The patent extracts and separately encodes phase spectral information from speech segments, distinguishing it from amplitude information. By isolating the phase component and applying targeted processing (such as phase unwrapping and selective encoding), the system maintains speech quality while reducing the amount of data that needs to be stored in the database.

Inventive Principle:
Principle #2Taking out (Extraction)

Solution Approach 2:

The patent changes the representation parameters of speech segments by encoding phase information in a compressed manner. Instead of storing complete high-resolution spectral data, the system uses parametric models to represent phase relationships, significantly reducing database footprint while preserving the auditory characteristics necessary for high-quality synthesis.

Inventive Principle:
Principle #35Parameter changes

2Reliability

If phase spectral information is encoded for all speech segments, then speech quality is improved, but encoding complexity increases

Engineering Contradiction:
Improvespeech qualityVSAvoidencoding complexity
Core Design Contradiction:
ReliabilityVSDevice complexity

Solution Approach 1:

The patent applies different encoding strategies to different portions of the speech signal based on local characteristics. Phase spectral information is encoded selectively for voiced segments where it provides the most benefit, while unvoiced segments use simpler encoding. This localized approach improves speech quality where needed without uniformly increasing encoding complexity across all segments.

Inventive Principle:
Principle #3Local quality

Solution Approach 2:

Instead of applying complex phase encoding to all speech segments, the patent applies it partially—specifically to voiced segments where phase information is most critical for quality. This selective application achieves the quality improvement goal with reduced overall encoding complexity compared to a universal approach.

Inventive Principle:
Principle #16Partial or excessive action

Data Source

PatentUS8280724B2Speech synthesis using complex spectral modeling
Publication Date: 2012.10.02 CERENCE OPERATING CO
  • US8280724B2 patent drawing
  • US8280724B2 patent drawing
  • US8280724B2 patent drawing

AI summary

A method for processing a speech signal includes dividing the speech signal into a succession of frames, identifying one or more of the frames as click frames, and extracting phase information from the click frames. The speech signal is encoded using the phase information. Methods are also provided for modeling phase spectra of voiced frames and click frames.