Synthesized Singing Voice Waveform Generation via Sub-Phonemic Dissection

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current text-to-speech synthesis systems lack the ability to generate realistic, human-like singing voices with emotional expression and flexible pitch control, which is essential for providing an expressive and engaging synthesized voice experience.

Innovation Solution

A computer program that receives song lyrics and a digital melody file, dissects the text and melody into sub-phonemic units and musical scores, matches these units with statistically trained contextual parametric models, and combines them with duration times to create a synthesized singing voice waveform, using a database of statistically trained models to represent the sound and pitch of each unit.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Reliability

If text-to-speech synthesis systems use conventional methods, then basic speech synthesis is achieved, but realistic human-like singing voice with emotional expression and flexible pitch control cannot be generated

Engineering Contradiction:
Improverealism of synthesized voiceVSAvoidpitch control flexibility
Core Design Contradiction:
ReliabilityVSAdaptability or versatility

Solution Approach 1:

The patent segments the singing voice synthesis process into distinct components: lyric text analysis, melody note extraction, phoneme-to-note mapping, and spectral envelope generation. Each component is handled by specialized modules that process specific aspects independently, allowing for realistic voice generation while maintaining flexible pitch control through the separate melody processing pipeline.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The system dynamically adjusts pitch and temporal parameters by mapping musical notes to phonemes with variable pitch contours and duration. The fundamental frequency (F0) is dynamically controlled to match the melody notes, enabling flexible pitch expression while maintaining natural speech-like transitions through statistical parametric modeling.

Inventive Principle:
Principle #15Dynamics

2Measurement precision

If the system dissects lyrics and melody into sub-phonemic units and musical scores with detailed processing, then accurate pitch and temporal control is achieved, but system complexity increases

Engineering Contradiction:
Improvepitch accuracyVSAvoidprocessing complexity
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The patent performs preliminary processing by pre-segmenting lyrics into phonemes and pre-extracting musical notes from the melody file before the actual synthesis. This preliminary action organizes the input data into structured formats (phoneme sequences, note sequences with pitch and duration) that simplify subsequent mapping and synthesis operations, reducing overall processing complexity while maintaining precision.

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The system introduces intermediate representations as mediators: phoneme sequences serve as intermediaries between text and speech output, while musical note sequences with extracted pitch and duration act as intermediaries between the melody file and the synthesized voice. These intermediaries structure the data in a way that simplifies the complex mapping relationships while preserving measurement precision.

Inventive Principle:
Principle #24Intermediary (Mediator)

3Ease of operation

If statistically trained contextual parametric models are used to represent sound of each sub-phonemic unit, then natural and expressive voice output is generated, but computational resources and processing time increase

Engineering Contradiction:
Improvenaturalness of synthesized voiceVSAvoidprocessing time
Core Design Contradiction:
Ease of operationVSLoss of time

Solution Approach 1:

The patent employs preliminary statistical training to create contextual parametric models that capture the acoustic characteristics of phonemes in different contexts. These pre-trained models are stored and can be rapidly applied during synthesis without requiring intensive computation at runtime, enabling natural and expressive voice output while reducing processing time during actual use.

Inventive Principle:
Principle #10Preliminary action

Data Source

PatentUS7977562B2Synthesized singing voice waveform generator
Publication Date: 2011.07.12 MICROSOFT TECHNOLOGY LICENSING LLC
  • US7977562B2 patent drawing
  • US7977562B2 patent drawing
  • US7977562B2 patent drawing

AI summary

Various technologies for generating a synthesized singing voice waveform. In one implementation, the computer program may receive a request from a user to create a synthesized singing voice using the lyrics of a song and a digital file containing its melody as inputs. The computer program may then dissect the lyrics' text and its melody file into its corresponding sub-phonemic units and musical score respectively. The musical score may be further dissected into a sequence of musical notes and duration times for each musical note. The computer program may then determine a fundamental frequency (F0), or pitch, of each musical note.