Synthesized Singing Voice Waveform Generation via Sub-Phonemic Dissection
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current text-to-speech synthesis systems lack the ability to generate realistic, human-like singing voices with emotional expression and flexible pitch control, which is essential for providing an expressive and engaging synthesized voice experience.
Innovation Solution
A computer program that receives song lyrics and a digital melody file, dissects the text and melody into sub-phonemic units and musical scores, matches these units with statistically trained contextual parametric models, and combines them with duration times to create a synthesized singing voice waveform, using a database of statistically trained models to represent the sound and pitch of each unit.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Reliability
If text-to-speech synthesis systems use conventional methods, then basic speech synthesis is achieved, but realistic human-like singing voice with emotional expression and flexible pitch control cannot be generated
Solution Approach 1:
The patent segments the singing voice synthesis process into distinct components: lyric text analysis, melody note extraction, phoneme-to-note mapping, and spectral envelope generation. Each component is handled by specialized modules that process specific aspects independently, allowing for realistic voice generation while maintaining flexible pitch control through the separate melody processing pipeline.
Solution Approach 2:
The system dynamically adjusts pitch and temporal parameters by mapping musical notes to phonemes with variable pitch contours and duration. The fundamental frequency (F0) is dynamically controlled to match the melody notes, enabling flexible pitch expression while maintaining natural speech-like transitions through statistical parametric modeling.
2Measurement precision
If the system dissects lyrics and melody into sub-phonemic units and musical scores with detailed processing, then accurate pitch and temporal control is achieved, but system complexity increases
Solution Approach 1:
The patent performs preliminary processing by pre-segmenting lyrics into phonemes and pre-extracting musical notes from the melody file before the actual synthesis. This preliminary action organizes the input data into structured formats (phoneme sequences, note sequences with pitch and duration) that simplify subsequent mapping and synthesis operations, reducing overall processing complexity while maintaining precision.
Solution Approach 2:
The system introduces intermediate representations as mediators: phoneme sequences serve as intermediaries between text and speech output, while musical note sequences with extracted pitch and duration act as intermediaries between the melody file and the synthesized voice. These intermediaries structure the data in a way that simplifies the complex mapping relationships while preserving measurement precision.
3Ease of operation
If statistically trained contextual parametric models are used to represent sound of each sub-phonemic unit, then natural and expressive voice output is generated, but computational resources and processing time increase
Solution Approach 1:
The patent employs preliminary statistical training to create contextual parametric models that capture the acoustic characteristics of phonemes in different contexts. These pre-trained models are stored and can be rapidly applied during synthesis without requiring intensive computation at runtime, enabling natural and expressive voice output while reducing processing time during actual use.
Data Source
AI summary
Various technologies for generating a synthesized singing voice waveform. In one implementation, the computer program may receive a request from a user to create a synthesized singing voice using the lyrics of a song and a digital file containing its melody as inputs. The computer program may then dissect the lyrics' text and its melody file into its corresponding sub-phonemic units and musical score respectively. The musical score may be further dissected into a sequence of musical notes and duration times for each musical note. The computer program may then determine a fundamental frequency (F0), or pitch, of each musical note.


