Electronic Musical Instrument Singing Voice Synthesis
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing electronic musical instruments using concatenative synthesis require extensive recording and complex calculations to produce a natural-sounding singing voice, making it time-consuming and inefficient.
Innovation Solution
An electronic musical instrument equipped with a trained acoustic model that performs machine learning on musical score data, lyric data, and singing voice data, allowing it to infer and synthesize a singing voice based on user input, eliminating the need for real-time singing and microphone usage.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Ease of manufacture
If concatenative synthesis is used to generate singing voice, then a singing voice can be produced by connecting fragments of recorded speech, but long hours of recording and complex calculations are required
Solution Approach 1:
The patent uses a trained acoustic model that has learned the singing voice characteristics of a singer from training data. Instead of requiring new recordings for each song, the model copies and reproduces the singer's voice characteristics through machine learning, eliminating the need for extensive new recording sessions
Solution Approach 2:
The acoustic model is trained in advance using training musical score data, training lyric data, and training singing voice data. This preliminary training action creates a reusable model that can generate singing voices without requiring time-consuming recording and processing for each new song
2Reliability
If concatenative synthesis is used to generate singing voice, then singing voice fragments can be connected, but complex calculations are required for smoothly joining fragments
Solution Approach 1:
The patent replaces the mechanical process of manually connecting speech fragments with complex calculations with a trained acoustic model that uses machine learning. The model automatically generates natural-sounding singing voices by learning from training data, substituting the complex fragment-joining mechanism with a learned predictive model
Solution Approach 2:
The acoustic model learns and reproduces singing voice characteristics by analyzing parameters from training data including pitch, timing, and spectral features. The model changes its internal parameters through training to accurately reproduce the singer's voice characteristics without requiring complex real-time fragment manipulation
3Productivity
If a trained acoustic model is used to generate singing voice, then recording time is reduced, but machine learning training is required
Solution Approach 1:
The acoustic model training is performed as a preliminary action before actual singing voice generation. The model learns from training musical score data, training lyric data, and training singing voice data in advance, creating a reusable asset that enables efficient real-time singing voice synthesis without requiring repeated training
Solution Approach 2:
The trained model copies the singer's voice characteristics into a reusable digital representation. This copied knowledge allows the system to generate singing voices efficiently for multiple songs without requiring the singer to re-record or re-train, transferring the complexity from ongoing operations to a one-time training process
Data Source
AI summary
An electronic musical instrument includes an operation unit that receives a user performance; and at least one processor. wherein the at least one processor performs the following: in accordance with a user operation specifying a chord on the operation unit, obtaining lyric data of a lyric and obtaining a plurality of pieces of waveform data respectively corresponding to a plurality of pitches indicated by the specified chord; inputting the obtained lyric data to a trained model that has been trained and learned singing voices of a singer so as to cause the trained model to output acoustic feature data in response thereto; synthesizing each of the plurality of pieces of waveform data with the acoustic feature data so as to generate a plurality of pieces of synthesized waveform data; and outputting a polyphonic synthesized singing voice based on the generated plurality of pieces of synthesized waveform data.


