Prosodically Modified Speech Unit Database for Natural Synthesis
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Speech unit selection synthesis often fails to produce consistently high-quality audio output due to limitations in the size and quality of speech databases, particularly when generating out-of-domain text, and lacks prosody modification, resulting in unsatisfactory prosodic contours.
Innovation Solution
A system that identifies and modifies speech units to match a desired prosodic curve by using signal processing techniques like Residual-Excited Linear Prediction (RELP) and Pitch Synchronous Overlap and Add (PSOLA), augmenting the speech unit database with transformed data to enhance prosodic coverage and quality.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Manufacturing precision
If domain-specific databases of speech samples are used to improve quality, then in-domain text produces high-quality speech, but out-of-domain text produces poor quality speech
Solution Approach 1:
The system transforms existing speech units by modifying their prosodic parameters (pitch, duration, energy) to create new variations. This allows the database to adapt to different domains and prosodic requirements without requiring separate domain-specific recordings, resolving the contradiction between speech quality and domain adaptability
Solution Approach 2:
The system segments speech into smaller units (phonemes, syllables, words) and recombines them with modified prosody. This segmentation allows flexible reconstruction of speech units with desired prosodic characteristics, enabling both high quality and domain versatility
2Reliability
If speech unit selection synthesis is used to generate natural audio output, then speech can be synthesized from selected units, but the quality is inconsistent due to database limitations
Solution Approach 1:
The system performs preliminary prosody modification on speech units during database construction, creating pre-modified units with various prosodic characteristics. This preliminary action ensures that when synthesis is performed, high-quality units with appropriate prosody are already available, improving consistency without requiring larger databases
Solution Approach 2:
The system creates modified copies of existing speech units with transformed prosody. These copies are added to the database, effectively multiplying the utility of original recordings and improving database quality consistency without proportionally increasing database size
3Quantity of substance
If previous techniques focus on segmental level and repurposing data to boost database size, then effective database size increases, but prosodic coverage remains limited
Solution Approach 1:
The system applies prosody modification techniques that transform the temporal and spectral characteristics of speech units, creating new prosodic patterns from existing data. This approach simultaneously increases effective database size and expands prosodic coverage by generating diverse prosodic variations from limited original recordings
Data Source
AI summary
Systems, methods, and computer-readable storage devices to improve the quality of synthetic speech generation. A system selects speech units from a speech unit database, the speech units corresponding to text to be converted to speech. The system identifies a desired prosodic curve of speech produced from the selected speech units, and also identifies an actual prosodic curve of the speech units. The selected speech units are modified such that a new prosodic curve of the modified speech units matches the desired prosodic curve. The system stores the modified speech units into the speech unit database for use in generating future speech, thereby increasing the prosodic coverage of the database with the expectation of improving the output quality.


