Speech Intonation Generation Using F0 Pattern Database
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Conventional speech synthesis technologies face challenges in generating natural intonation, particularly in controlling accent parameters and expressing dynamic speech characteristics, leading to homogenized intonation and limitations in handling unregistered words and varying sentence patterns.
Innovation Solution
A method for generating intonation patterns in speech synthesis that estimates an outline based on language information, selects an intonation pattern from a database of actual speech, and adjusts frequency levels, allowing for flexible and accurate reproduction of speaker characteristics without relying on prosodic categories.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Adaptability or versatility
If conventional generation models (Fujisaki Model) are used to generate intonation patterns, then the model can flexibly express intensities and positions of accents, but the parameters become excessively simplified and cannot precisely control speaker characteristics and speech styles
Solution Approach 1:
The patent copies actual F0 patterns from recorded speech data and stores them in a database. Instead of generating intonation patterns through complex mathematical models, the system directly uses recorded F0 patterns as templates, thereby preserving the natural precision of human speech characteristics while simplifying the generation process.
Solution Approach 2:
The patent replaces the mechanical/mathematical model-based intonation generation system with a data-driven approach using actual recorded F0 patterns. This substitution eliminates the need for complex parameter calculations and directly uses empirical speech data, thereby improving precision without increasing computational complexity.
2Ease of manufacture
If F0 patterns are selected based on prosodic categories from language information, then the selection process is systematic, but appropriate F0 patterns cannot be applied when text cannot be classified into existing prosodic categories
Solution Approach 1:
The patent makes the F0 pattern selection process dynamic by allowing direct matching of F0 patterns based on similarity metrics rather than fixed prosodic category classification. This dynamic approach enables the system to adapt to any text input by finding the most similar F0 pattern in the database, regardless of whether the text fits into predefined prosodic categories.
Solution Approach 2:
The patent creates a universal F0 pattern database that can serve multiple functions: it can be queried using prosodic categories when available, but more importantly, it can directly match any text input based on F0 pattern similarity. This universal database structure handles both classified and unclassified text equally well.
3Ease of manufacture
If F0 patterns are equated and modeled to create representative patterns, then the processing becomes simpler, but F0 variations in the database cannot be sufficiently expressed
Solution Approach 1:
The patent segments the F0 pattern database into individual, distinct F0 pattern records rather than creating a single averaged model. Each F0 pattern from the recorded speech database is preserved as a separate, authentic template, thereby maintaining the diversity and richness of original F0 variations while keeping each individual pattern simple and easy to process.
Data Source
AI summary
In generation of an intonation pattern of a speech synthesis, a speech synthesis system is capable of providing a highly natural speech and capable of reproducing speech characteristics of a speaker flexibly and accurately by effectively utilizing FO patterns of actual speech accumulated in a database. An intonation generation method generates an intonation of synthesized speech for text by estimating, based on language information of the text and based on the estimated outline of the intonation, and then selects an optimum intonation pattern from a database which stores intonation patterns of actual speech. Speech characteristics recorded in advance are reflected in an estimation of an outline of the intonation pattern and selection of a waveform element of a speech.


