Prosody Generator Using Density-Based Method Selection
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current methods for generating prosody information in speech synthesis, whether using rules or statistical techniques, face challenges in producing natural-sounding speech, particularly with sparse data, leading to unnatural sound and displaced accent positions, and require large datasets that are difficult to collect.
Innovation Solution
A prosody generator that divides the data space into subspaces and extracts density information to selectively choose between statistical and rule-based methods for prosody information generation, using statistical techniques for dense areas and rule-based methods for sparse areas, thereby avoiding the need for extensive data collection.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Reliability
If statistical techniques are used for prosody generation, then human-like prosody information is generated, but sparse data portions cause incorrect F0 patterns and disturbed accent positions
Solution Approach 1:
The learning data space is divided into multiple subspaces based on feature quantity characteristics. Each subspace is processed independently with appropriate method selection, allowing dense subspaces to use statistical techniques while sparse subspaces use rule-based methods, thereby resolving the contradiction between data quantity requirements and accuracy.
Solution Approach 2:
Different prosody generation methods are applied to different subspaces based on their local data density characteristics. Dense subspaces receive statistical processing while sparse subspaces receive rule-based processing, ensuring each region gets the most appropriate treatment for its specific conditions rather than applying a uniform approach.
2Ease of manufacture
If rule-based methods are used for prosody generation, then simple F0 pattern generation is achieved, but the synthesized speech sounds mechanical and unnatural
Solution Approach 1:
The prosody generation process is segmented into different method applications: rule-based methods handle simple, sparse cases while statistical methods handle complex, dense cases. This segmentation allows the system to leverage the simplicity of rules where appropriate while using statistics where naturalness is critical.
Solution Approach 2:
The system dynamically selects between rule-based and statistical methods based on the characteristics of each subspace. This dynamic adaptation allows the system to optimize between simplicity and naturalness in real-time, choosing the most appropriate method for each specific situation rather than being fixed to one approach.
3Adaptability or versatility
If data space is divided into clusters based on information quantity, then statistical processing becomes feasible, but sparse and dense portions are created causing inconsistent prosody quality
Solution Approach 1:
The system applies different processing qualities to different subspaces: statistical processing for dense subspaces and rule-based processing for sparse subspaces. This local differentiation maintains consistency within each subspace while accommodating the varying data densities across the entire data space.
Solution Approach 2:
By segmenting the data space and applying appropriate methods to each segment, the system maintains stable, consistent prosody quality within each subspace. The segmentation prevents the inconsistency that would arise from applying a single method uniformly across subspaces with vastly different data densities.
Data Source
AI summary
There is provided a prosody generator that generates prosody information for implementing highly natural speech synthesis without unnecessarily collecting large quantities of learning data. A data dividing means 81 divides into subspaces the data space of a learning database as an assembly of learning data indicative of the feature quantities of speech waveforms. A density information extracting means 82 extracts density information indicative of the density state in terms of information quantity of the learning data in each of the subspaces divided by the data dividing means 81. A prosody information generating method selecting means 83 selects either a first method or a second method as a prosody information generating method based on the density information, the first method involving generating the prosody information using a statistical technique, the second method involving generating the prosody information using rules based on heuristics.


