Singing Synthesis Database Creation Using Melody Component Extraction
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Conventional voice synthesis techniques using Hidden Markov Models (HMMs) struggle to accurately model the unique singing expression and melody style of a singing person, as they are based on phonemes and do not effectively capture variation in fundamental frequency independent of phoneme context.
Innovation Solution
A singing synthesizing database creation apparatus that extracts and models the melody component parameters from learning waveform data and score data, using machine learning to generate parameters that represent the variation in fundamental frequency between notes, allowing for accurate modeling and synthesis of singing voices that reflect a singing person's unique style.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Reliability
If voice synthesis uses HMM based on phonemes, then voice synthesis can be performed with compact database size, but singing expression and melody style cannot be accurately modeled
Solution Approach 1:
The patent segments the pitch variation into two independent components: phoneme-dependent pitch variation and melody-dependent pitch variation. This segmentation allows each component to be modeled separately using HMMs, enabling accurate representation of singing expression while maintaining the efficiency of phoneme-based synthesis. The melody component is extracted by analyzing pitch curves and isolating variations that occur independently of phoneme boundaries.
Solution Approach 2:
The patent introduces a new dimension of modeling by adding melody-specific HMM parameters alongside the existing phoneme-based HMM parameters. This dimensional expansion allows the system to capture pitch variations in the melody dimension that are independent of phoneme context, thereby accurately representing singing expression without abandoning the efficient phoneme-based framework.
2Loss of information
If phoneme-based HMM modeling is used, then database size is reduced, but variation in fundamental frequency independent of phoneme context cannot be captured
Solution Approach 1:
The patent segments pitch variation into phoneme-dependent and melody-dependent components, allowing selective storage of only melody-specific parameters. This segmentation enables the system to preserve essential singing expression information while avoiding redundant storage of phoneme-level details that are already captured by existing phoneme HMMs.
Solution Approach 2:
The patent merges the phoneme-based HMM framework with melody-specific HMM parameters to create a hybrid model. This combination allows the system to leverage the compactness of phoneme-based representation while incorporating additional melody style information, achieving both data efficiency and expressive accuracy.
3Productivity
If conventional HMM parameters are used, then synthesis efficiency is maintained, but naturalness of synthesized singing voices deteriorates
Solution Approach 1:
The patent segments the synthesis process into phoneme-level synthesis (maintaining efficiency) and melody-level pitch adjustment (improving naturalness). By separating these functions, the system can use efficient phoneme-based segment connection while applying melody-specific pitch curves to achieve natural singing expression.
Solution Approach 2:
The patent performs preliminary extraction and modeling of melody components from training data before the actual synthesis process. This preliminary action creates pre-computed melody HMM parameters that can be efficiently applied during synthesis, maintaining productivity while ensuring naturalness through accurate pre-modeled singing style characteristics.
Data Source
AI summary
Waveform data representative of singing voices of a singing music piece are analyzed to generate melody component data representative of variation over time in fundamental frequency component presumed to represent a melody in the singing voices. Then, through machine learning that uses score data representative of a musical score of the singing music piece and the melody component data, a melody component model, representative of a variation component presumed to represent the melody among the variation over time in fundamental frequency component, is generated for each combination of notes. Parameters defining the melody component models and note identifiers indicative of the combinations of notes whose variation over time in fundamental frequency component are represented by the melody component models are stored into a pitch curve generating database in association with each other.


