Electronic Musical Instrument Controller for Natural Singing Voice Synthesis

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing electronic musical instruments struggle to produce a natural singing voice due to the joining of waveform data pieces for syllable-by-syllable pronunciation and loop playback, which requires large memory capacity and results in an unnatural sound.

Innovation Solution

An information processing device with a controller that continues sound emission of a vowel based on a parameter for a vowel frame until the operation on the operation element is released, even if the operation continues after the start of sound emission of a syllable.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Manufacturing precision

If waveform data pieces of utterance units are joined together for syllable-by-syllable pronunciation and loop playback, then syllable-by-syllable pronunciation is achieved, but the sound becomes unnatural and large memory capacity is required

Engineering Contradiction:
Improvepronunciation accuracyVSAvoidsound naturalness
Core Design Contradiction:
Manufacturing precisionVSReliability

Solution Approach 1:

The patent uses a singing voice parameter sequence that copies the temporal and spectral characteristics of actual human singing voices. Instead of joining discrete waveform pieces, the system generates continuous sound by applying a sequence of parameters (pitch, timbre, intensity) that mirror natural singing patterns, thereby achieving both pronunciation accuracy and natural sound quality

Inventive Principle:
Principle #26Copying

Solution Approach 2:

The patent implements dynamic parameter adjustment throughout the sound generation process. The singing voice parameters vary continuously over time to match the natural dynamics of human singing, including pitch contours, timbre changes, and intensity variations. This dynamic approach replaces the static joining of waveform pieces with a flexible, adaptive parameter-based system that produces natural-sounding results

Inventive Principle:
Principle #15Dynamics

2Manufacturing precision

If waveform data pieces of utterance units are chronologically sequenced, then syllable-by-syllable pronunciation is achieved, but large memory capacity is required

Engineering Contradiction:
Improvepronunciation accuracyVSAvoidmemory capacity
Core Design Contradiction:
Manufacturing precisionVSQuantity of substance

Solution Approach 1:

Instead of storing large waveform data pieces for each utterance unit, the patent copies only the essential singing voice parameters (pitch, timbre, intensity contours) that define the sound characteristics. These compact parameter sequences require minimal memory while still enabling accurate syllable-by-syllable pronunciation when applied to a sound generator

Inventive Principle:
Principle #26Copying

Solution Approach 2:

The patent transforms the representation of vocal sounds from storing complete waveform data to storing and manipulating a small set of controlling parameters. By changing from waveform-based storage to parameter-based representation, the system achieves the same pronunciation accuracy with dramatically reduced memory requirements, as parameters can be mathematically applied to generate the actual sound waves

Inventive Principle:
Principle #35Parameter changes

Data Source

PatentUS20250111844A1Information processing device, electronic musical instrument, electronic musical instrument system, method, and storage medium
Publication Date: 2025.04.03 CASIO COMPUTER CO LTD
  • US20250111844A1 patent drawing
  • US20250111844A1 patent drawing
  • US20250111844A1 patent drawing

AI summary

An information processing device includes a controller. In response to detection of an operation on an operator, the controller causes sound emission of a syllable to start based on a parameter for a syllable start frame. In a case where the operation continues even after start of sound emission of a vowel based on a parameter for a vowel frame in a vowel section included in the syllable, the controller causes the sound emission of the vowel based on the parameter for the vowel frame to continue until the operation is released.