Singing Synthesis Database Creation Using Melody Component Extraction

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Conventional voice synthesis techniques using Hidden Markov Models (HMMs) struggle to accurately model the unique singing expression and melody style of a singing person, as they are based on phonemes and do not effectively capture variation in fundamental frequency independent of phoneme context.

Innovation Solution

A singing synthesizing database creation apparatus that extracts and models the melody component parameters from learning waveform data and score data, using machine learning to generate parameters that represent the variation in fundamental frequency between notes, allowing for accurate modeling and synthesis of singing voices that reflect a singing person's unique style.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Reliability

If voice synthesis uses HMM based on phonemes, then voice synthesis can be performed with compact database size, but singing expression and melody style cannot be accurately modeled

Engineering Contradiction:
Improveaccuracy of singing expression modelingVSAvoidcomplexity of modeling system
Core Design Contradiction:
ReliabilityVSDevice complexity

Solution Approach 1:

The patent segments the pitch variation into two independent components: phoneme-dependent pitch variation and melody-dependent pitch variation. This segmentation allows each component to be modeled separately using HMMs, enabling accurate representation of singing expression while maintaining the efficiency of phoneme-based synthesis. The melody component is extracted by analyzing pitch curves and isolating variations that occur independently of phoneme boundaries.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent introduces a new dimension of modeling by adding melody-specific HMM parameters alongside the existing phoneme-based HMM parameters. This dimensional expansion allows the system to capture pitch variations in the melody dimension that are independent of phoneme context, thereby accurately representing singing expression without abandoning the efficient phoneme-based framework.

Inventive Principle:
Principle #17Another dimension (Dimensionality change)

2Loss of information

If phoneme-based HMM modeling is used, then database size is reduced, but variation in fundamental frequency independent of phoneme context cannot be captured

Engineering Contradiction:
Improveloss of melody style informationVSAvoidamount of stored data
Core Design Contradiction:
Loss of informationVSQuantity of substance

Solution Approach 1:

The patent segments pitch variation into phoneme-dependent and melody-dependent components, allowing selective storage of only melody-specific parameters. This segmentation enables the system to preserve essential singing expression information while avoiding redundant storage of phoneme-level details that are already captured by existing phoneme HMMs.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent merges the phoneme-based HMM framework with melody-specific HMM parameters to create a hybrid model. This combination allows the system to leverage the compactness of phoneme-based representation while incorporating additional melody style information, achieving both data efficiency and expressive accuracy.

Inventive Principle:
Principle #5Merging (Combining)

3Productivity

If conventional HMM parameters are used, then synthesis efficiency is maintained, but naturalness of synthesized singing voices deteriorates

Engineering Contradiction:
Improvesynthesis efficiencyVSAvoidnaturalness of synthesized voice
Core Design Contradiction:
ProductivityVSReliability

Solution Approach 1:

The patent segments the synthesis process into phoneme-level synthesis (maintaining efficiency) and melody-level pitch adjustment (improving naturalness). By separating these functions, the system can use efficient phoneme-based segment connection while applying melody-specific pitch curves to achieve natural singing expression.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent performs preliminary extraction and modeling of melody components from training data before the actual synthesis process. This preliminary action creates pre-computed melody HMM parameters that can be efficiently applied during synthesis, maintaining productivity while ensuring naturalness through accurate pre-modeled singing style characteristics.

Inventive Principle:
Principle #10Preliminary action

Data Source

PatentUS8115089B2Apparatus and method for creating singing synthesizing database, and pitch curve generation apparatus and method
Publication Date: 2012.02.14 YAMAHA CORP
  • US8115089B2 patent drawing
  • US8115089B2 patent drawing
  • US8115089B2 patent drawing

AI summary

Waveform data representative of singing voices of a singing music piece are analyzed to generate melody component data representative of variation over time in fundamental frequency component presumed to represent a melody in the singing voices. Then, through machine learning that uses score data representative of a musical score of the singing music piece and the melody component data, a melody component model, representative of a variation component presumed to represent the melody among the variation over time in fundamental frequency component, is generated for each combination of notes. Parameters defining the melody component models and note identifiers indicative of the combinations of notes whose variation over time in fundamental frequency component are represented by the melody component models are stored into a pitch curve generating database in association with each other.