Electronic Musical Instrument Controller for Natural Singing Voice Synthesis
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing electronic musical instruments struggle to produce a natural singing voice due to the joining of waveform data pieces for syllable-by-syllable pronunciation and loop playback, which requires large memory capacity and results in an unnatural sound.
Innovation Solution
An information processing device with a controller that continues sound emission of a vowel based on a parameter for a vowel frame until the operation on the operation element is released, even if the operation continues after the start of sound emission of a syllable.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Manufacturing precision
If waveform data pieces of utterance units are joined together for syllable-by-syllable pronunciation and loop playback, then syllable-by-syllable pronunciation is achieved, but the sound becomes unnatural and large memory capacity is required
Solution Approach 1:
The patent uses a singing voice parameter sequence that copies the temporal and spectral characteristics of actual human singing voices. Instead of joining discrete waveform pieces, the system generates continuous sound by applying a sequence of parameters (pitch, timbre, intensity) that mirror natural singing patterns, thereby achieving both pronunciation accuracy and natural sound quality
Solution Approach 2:
The patent implements dynamic parameter adjustment throughout the sound generation process. The singing voice parameters vary continuously over time to match the natural dynamics of human singing, including pitch contours, timbre changes, and intensity variations. This dynamic approach replaces the static joining of waveform pieces with a flexible, adaptive parameter-based system that produces natural-sounding results
2Manufacturing precision
If waveform data pieces of utterance units are chronologically sequenced, then syllable-by-syllable pronunciation is achieved, but large memory capacity is required
Solution Approach 1:
Instead of storing large waveform data pieces for each utterance unit, the patent copies only the essential singing voice parameters (pitch, timbre, intensity contours) that define the sound characteristics. These compact parameter sequences require minimal memory while still enabling accurate syllable-by-syllable pronunciation when applied to a sound generator
Solution Approach 2:
The patent transforms the representation of vocal sounds from storing complete waveform data to storing and manipulating a small set of controlling parameters. By changing from waveform-based storage to parameter-based representation, the system achieves the same pronunciation accuracy with dramatically reduced memory requirements, as parameters can be mathematically applied to generate the actual sound waves
Data Source
AI summary
An information processing device includes a controller. In response to detection of an operation on an operator, the controller causes sound emission of a syllable to start based on a parameter for a syllable start frame. In a case where the operation continues even after start of sound emission of a vowel based on a parameter for a vowel frame in a vowel section included in the syllable, the controller causes the sound emission of the vowel based on the parameter for the vowel frame to continue until the operation is released.


