Hybrid Speech Synthesis System for Natural Audio and Storage Efficiency
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current speech synthesis methods either produce synthetic sounding speech efficiently with model-based approaches or natural speech but are inflexible and require large storage with segment-concatenation methods, lacking the ability to dynamically construct 'unseen' sounds.
Innovation Solution
A hybrid speech synthesis system that combines model-based and template-based approaches, using statistical speech models and psycho-acoustic rules to select and sequence speech segments, allowing for the generation of natural and flexible speech synthesis by optimally merging model and template segments.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Quantity of substance
If model-based speech synthesis is used, then storage efficiency and flexibility are improved, but speech quality becomes synthetic and processed
Solution Approach 1:
The patent combines model-based speech synthesis with template-based speech segments to create a hybrid system. The model generates speech parameters efficiently while templates provide natural speech quality references, merging the advantages of both approaches to resolve the contradiction between storage efficiency and speech quality.
Solution Approach 2:
The system uses composite speech representation by combining statistically modeled speech parameters with actual recorded speech templates. This composite approach allows the system to maintain storage efficiency through modeling while incorporating natural speech characteristics from templates to improve overall speech quality.
2Manufacturing precision
If segment-concatenation speech synthesis is used, then speech quality becomes natural, but storage requirements increase and flexibility decreases
Solution Approach 1:
The system extracts only the essential natural speech characteristics from recorded templates rather than storing complete speech segments. By extracting key acoustic features and parameters from natural speech while discarding redundant information, the system maintains speech quality while reducing storage requirements.
Solution Approach 2:
The patent segments speech into different representational components (statistical parameters, template features, acoustic models) that can be stored and processed separately. This segmentation allows the system to use compact statistical representations for most speech while incorporating template data only where needed for quality enhancement.
3Manufacturing precision
If segment-concatenation is used, then natural speech is produced, but the system becomes inflexible and cannot generate unseen sounds
Solution Approach 1:
The system dynamically switches between using statistical models and template-based approaches depending on the speech context and availability of template data. This dynamic adaptation allows the system to generate natural speech when templates are available while falling back to flexible model-based generation for unseen sounds, resolving the contradiction between naturalness and flexibility.
Solution Approach 2:
The statistical speech model acts as an intermediary between the template database and the final speech output. When template data is unavailable or insufficient, the model generates speech parameters that interpolate between existing templates, enabling the system to produce unseen sounds while maintaining natural speech characteristics through the mediating model.
Data Source
Figure 1
Figure 2
Figure 3
AI summary
A speech synthesis system and method is described. A speech segment database references speech segments having various different speech representational structures. A speech segment selector selects from the speech segment database a sequence of speech segment candidates corresponding to a target text. A speech segment sequencer generates from the speech segment candidates sequenced speech segments corresponding to the target text. A speech segment synthesizer combines the selected sequenced speech segments to produce a synthesized speech signal output corresponding to the target text.