Hybrid Speech Synthesis System for Natural Audio and Storage Efficiency

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current speech synthesis methods either produce synthetic sounding speech efficiently with model-based approaches or natural speech but are inflexible and require large storage with segment-concatenation methods, lacking the ability to dynamically construct 'unseen' sounds.

Innovation Solution

A hybrid speech synthesis system that combines model-based and template-based approaches, using statistical speech models and psycho-acoustic rules to select and sequence speech segments, allowing for the generation of natural and flexible speech synthesis by optimally merging model and template segments.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Quantity of substance

If model-based speech synthesis is used, then storage efficiency and flexibility are improved, but speech quality becomes synthetic and processed

Engineering Contradiction:
Improvestorage efficiencyVSAvoidspeech quality
Core Design Contradiction:
Quantity of substanceVSManufacturing precision

Solution Approach 1:

The patent combines model-based speech synthesis with template-based speech segments to create a hybrid system. The model generates speech parameters efficiently while templates provide natural speech quality references, merging the advantages of both approaches to resolve the contradiction between storage efficiency and speech quality.

Inventive Principle:
Principle #5Merging (Combining)

Solution Approach 2:

The system uses composite speech representation by combining statistically modeled speech parameters with actual recorded speech templates. This composite approach allows the system to maintain storage efficiency through modeling while incorporating natural speech characteristics from templates to improve overall speech quality.

Inventive Principle:
Principle #40Composite materials

2Manufacturing precision

If segment-concatenation speech synthesis is used, then speech quality becomes natural, but storage requirements increase and flexibility decreases

Engineering Contradiction:
Improvespeech qualityVSAvoidstorage requirements
Core Design Contradiction:
Manufacturing precisionVSQuantity of substance

Solution Approach 1:

The system extracts only the essential natural speech characteristics from recorded templates rather than storing complete speech segments. By extracting key acoustic features and parameters from natural speech while discarding redundant information, the system maintains speech quality while reducing storage requirements.

Inventive Principle:
Principle #2Taking out (Extraction)

Solution Approach 2:

The patent segments speech into different representational components (statistical parameters, template features, acoustic models) that can be stored and processed separately. This segmentation allows the system to use compact statistical representations for most speech while incorporating template data only where needed for quality enhancement.

Inventive Principle:
Principle #1Segmentation

3Manufacturing precision

If segment-concatenation is used, then natural speech is produced, but the system becomes inflexible and cannot generate unseen sounds

Engineering Contradiction:
Improvespeech naturalnessVSAvoidflexibility
Core Design Contradiction:
Manufacturing precisionVSAdaptability or versatility

Solution Approach 1:

The system dynamically switches between using statistical models and template-based approaches depending on the speech context and availability of template data. This dynamic adaptation allows the system to generate natural speech when templates are available while falling back to flexible model-based generation for unseen sounds, resolving the contradiction between naturalness and flexibility.

Inventive Principle:
Principle #15Dynamics

Solution Approach 2:

The statistical speech model acts as an intermediary between the template database and the final speech output. When template data is unavailable or insufficient, the model generates speech parameters that interpolate between existing templates, enabling the system to produce unseen sounds while maintaining natural speech characteristics through the mediating model.

Inventive Principle:
Principle #24Intermediary (Mediator)

Data Source

PatentEP2179414B1Synthesis by generation and concatenation of multi-form segments
Publication Date: 2018.11.07 NUANCE COMMUNICATIONS INC
  • EP2179414B1 patent drawingFigure 1
  • EP2179414B1 patent drawingFigure 2
  • EP2179414B1 patent drawingFigure 3

AI summary

A speech synthesis system and method is described. A speech segment database references speech segments having various different speech representational structures. A speech segment selector selects from the speech segment database a sequence of speech segment candidates corresponding to a target text. A speech segment sequencer generates from the speech segment candidates sequenced speech segments corresponding to the target text. A speech segment synthesizer combines the selected sequenced speech segments to produce a synthesized speech signal output corresponding to the target text.