Electronic musical instrument, lyrics progression control method and computer program product
Through a multi-key structure and a deep-learning acoustic model, the electronic musical instrument achieves flexible lyric progression control, solving the problem of high difficulty in lyric progression operation in existing technologies, reducing costs and improving the expressiveness of synthesized sound.
Patent Information
- Application Number
- CN202111041397.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Priority Date
- 2020-09-08
- Filing Date
- 2021-09-07
- Publication Date
- 2025-09-19
- Estimated Expiration
- 2041-09-07
AI Technical Summary
Existing electronic musical instruments have a high threshold for controlling the progression of lyrics, making it difficult to easily pronounce lyrics with synthesized sound, and user operation is not flexible enough.
An electronic musical instrument with a multi-key structure uses a processor to control the position of syllables and the pronunciation of notes in a phrase through key operations in the first and second ranges, and combines statistical sound synthesis technology and acoustic models based on deep learning to achieve automatic control of the progression of lyrics.
It achieves more flexible lyrics progression control, reduces user operation difficulty, improves the expressiveness of synthesized sound, reduces storage requirements and production time, and reduces the cost of electronic musical instruments.
Smart Images

Figure CN114155822B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to an electronic musical instrument, a method and a program. Background Art
[0002] In recent years, the use of synthesized sounds has increased. Among them, if there is an electronic musical instrument that can not only automatically play but also output synthesized sounds corresponding to the lyrics in response to the user's (player's) keystrokes, it will be more flexible in expressing synthesized sounds and is therefore preferred.
[0003] Operating the progress of a musical phrase (eg, lyrics) related to a performance using a dedicated controller has a high threshold from the user's operation point of view, and it is difficult to easily appreciate the pronunciation of the lyrics using synthesized sounds. Summary of the Invention
[0004] One object of the present invention is to provide an electronic musical instrument, method, and program that can appropriately control the progression of a musical phrase (eg, lyrics) related to a performance.
[0005] An electronic musical instrument according to one technical solution of the present invention comprises: a plurality of keys, including a plurality of first keys corresponding to a first musical range and a plurality of second keys corresponding to a second musical range; and at least one processor; the at least one processor determines the position of a syllable included in a musical phrase based on key operations in the first musical range; and the at least one processor instructs the pronunciation of a sound corresponding to the determined syllable position based on key operations in the second musical range. BRIEF DESCRIPTION OF THE DRAWINGS
[0006] Figure 1 1 is a diagram showing an example of the appearance of an electronic musical instrument 10 according to one embodiment.
[0007] Figure 2 1 is a diagram showing an example of the hardware configuration of a control system 200 of an electronic musical instrument 10 according to an embodiment.
[0008] Figure 3 1 is a diagram showing a configuration example of the voice learning unit 301 according to one embodiment.
[0009] Figure 4 This is a diagram showing an example of the waveform data output unit 211 according to one embodiment.
[0010] Figure 5 This is a diagram showing another example of the waveform data output unit 211 according to one embodiment.
[0011] Figure 6 This is a diagram showing an example of a keyboard key range division for syllable position control according to one embodiment.
[0012] Figures 7A to 7C This is a diagram showing an example of syllables assigned to the control key range.
[0013] Figure 8 This is a diagram showing an example of a flowchart of a lyrics progression control method according to one embodiment.
[0014] Figure 9 This is a diagram showing an example of a flowchart of a syllable position control process according to one embodiment.
[0015] Figure 10 This is a diagram showing an example of a flowchart of a performance control process according to one embodiment.
[0016] Figure 11 This is a diagram showing an example of a flowchart of syllable progression discrimination processing according to one embodiment.
[0017] Figure 12 This is a diagram showing an example of a flowchart of a syllable change process according to one embodiment.
[0018] Figure 13A and Figure 13B This is a diagram showing an example of the appearance of the keys of the control key range.
[0019] Figure 14 This is a diagram showing an example of a tablet terminal that implements a lyrics progression control method according to an embodiment. DETAILED DESCRIPTION
[0020] Hereinafter, embodiments of the present invention will be described in detail with reference to the accompanying drawings. In the following description, identical parts are given identical reference numerals. Since the names and functions of identical parts are identical, detailed descriptions thereof will not be repeated.
[0021] (Electronic Musical Instruments)
[0022] Figure 1 140 ] is a diagram showing an example of the appearance of an electronic musical instrument 10 according to one embodiment. The electronic musical instrument 10 may be equipped with a switch (button) panel 140b, a keyboard 140k, pedals 140p, a display 150d, a speaker 150s, and the like.
[0023] The electronic musical instrument 10 is a device that receives user input via operating elements such as a keyboard and switches, and controls performance and lyric progression. The electronic musical instrument 10 can be a device that generates sound corresponding to performance information such as MIDI (Musical Instrument Digital Interface) data. This device can be an electronic musical instrument (such as an electronic piano or a synthesizer) or a simulated musical instrument equipped with sensors and configured to have the functions of the aforementioned operating elements.
[0024] The switch panel 140b may include switches for operating volume designation, sound source, timbre, etc., song (accompaniment) selection, song playback start / stop, song playback setting (beat, etc.), etc.
[0025] The keyboard 140k may have a plurality of keys as performance operating elements. The pedal 140p may be a sustain pedal that prolongs the sound of a depressed key while the pedal is depressed, or a pedal for operating a controller that processes the tone, volume, etc.
[0026] In the present invention, the damper pedal, pedal, foot switch, controller (operating element), switch, button, touch panel, etc. can be interchangeably referred to. Depressing the pedal in the present invention can be interchangeably referred to as operating the controller.
[0027] The keys may also be referred to as playing operators, pitch operators, tone operators, direct operators, first operators, etc. The pedals may also be referred to as non-playing operators, non-pitch operators, non-tone operators, indirect operators, second operators, etc.
[0028] The display 150d can display lyrics, musical scores, various setting information, etc. The speaker 150s can be used to emit sounds generated by the performance.
[0029] Furthermore, the electronic musical instrument 10 may be capable of generating or converting at least one of a MIDI message (event) and an OSC (Open Sound Control) message.
[0030] The electronic musical instrument 10 may also be referred to as a control device 10 , a syllable progression control device 10 , or the like.
[0031] The electronic musical instrument 10 can communicate with a network (such as the Internet) via at least one of wired and wireless communication (for example, LTE (Long Term Evolution), 5G NR (5th generation mobile communication system New Radio), Wi-Fi (registered trademark), etc.).
[0032] The electronic musical instrument 10 can store in advance, transmit, and / or receive vocal data (also referred to as lyric text data, lyric information, etc.) related to lyrics for the musical process control target via a network. The vocal data can be text written in a musical notation language (e.g., MusicXML), expressed in a MIDI data storage format (e.g., SMF (Standard MIDI File) format), or provided as a conventional text file. The vocal data can be the vocal data 215 described below. In the present invention, vocals, sounds, and the like are interchangeable.
[0033] Furthermore, the electronic musical instrument 10 can obtain the content of the user singing in real time via a microphone or the like included in the electronic musical instrument 10 , and obtain text data obtained by applying voice recognition processing to the content as singing voice data.
[0034] Figure 2 1 is a diagram showing an example of the hardware configuration of a control system 200 of an electronic musical instrument 10 according to an embodiment.
[0035] Central Processing Unit (CPU) 201, ROM (Read Only Memory) 202, RAM (Random Access Memory) 203, Waveform Data Output Unit 211, Figure 1 The switch (button) panel 140b, the keyboard 140k, the key scanner 206 connected to the pedal 140p, and the Figure 1 An LCD controller 208 connected to an LCD (Liquid Crystal Display) as an example of the display 150 d is connected to the system bus 209 .
[0036] A timer 210 (also called a counter) for controlling the performance may also be connected to the CPU 201. The timer 210 may be used, for example, to count the number of times the electronic musical instrument 10 automatically performs the performance. The CPU 201 may also be called a processor, and may include an interface with peripheral circuits, a control circuit, an arithmetic circuit, registers, and the like.
[0037] CPU 201 uses RAM 203 as a working memory while executing the control program stored in ROM 202, thereby executing Figure 1 The ROM 202 may store vocal data, accompaniment data, song data including these data, and the like in addition to the control program and various fixed data.
[0038] The waveform data output unit 211 may include a sound source LSI (Large Scale Integrated Circuit) 204, a sound synthesis LSI 205, etc. The sound source LSI 204 and the sound synthesis LSI 205 may also be integrated into one LSI. Figure 3 In addition, part of the processing of the waveform data output unit 211 may be performed by the CPU 201 or by a CPU included in the waveform data output unit 211 .
[0039] The singing voice waveform data 217 and song waveform data 218 output from the waveform data output unit 211 are converted into analog singing voice output signals and analog musical sound output signals by D / A converters 212 and 213, respectively. The analog musical sound output signal and the analog singing voice output signal may be mixed by a mixer 214, and the mixed signal amplified by an amplifier 215 is then output from the speaker 150s or an output terminal. Alternatively, the singing voice waveform data may be referred to as synthesized singing voice data. Although not shown, the singing voice waveform data 217 and song waveform data 218 may be digitally synthesized and then converted into analog values by a D / A converter to obtain a mixed signal.
[0040] The key scanner 206 always scans Figure 1 The key-press / key-release status of the keyboard 140k, the switch operation status of the switch panel 140b, the pedal operation status of the pedal 140p, etc. are interrupted to the CPU 201 to convey the status change.
[0041] The LCD controller 208 is an IC (Integrated Circuit) that controls the display state of an LCD as an example of the display 150 d .
[0042] The system configuration is merely an example and is not intended to be limiting. For example, the number of circuits included is not limited to this. The electronic musical instrument 10 may also have a configuration that excludes some circuits (mechanisms), or may have a configuration in which multiple circuits perform the functions of a single circuit. Alternatively, a single circuit may perform the functions of multiple circuits.
[0043] Furthermore, the electronic musical instrument 10 may include hardware such as a microprocessor, a digital signal processor (DSP), an ASIC (Application Specific Integrated Circuit), a PLD (Programmable Logic Device), or an FPGA (Field Programmable Gate Array), and may implement some or all of the functional blocks using this hardware. For example, the CPU 201 may be implemented using at least one of these hardware components.
[0044] <Generation of Acoustic Model>
[0045] Figure 3 The figure shows an example of the structure of the voice learning unit 301 according to one embodiment. The voice learning unit 301 can be used as Figure 1 The electronic musical instrument 10 is independently installed as a function executed by the external server computer 300. In addition, the sound learning unit 301 can also be built into the electronic musical instrument 10 as a function executed by the CPU 201, the sound synthesis LSI 205, etc.
[0046] The speech learning unit 301 and the waveform data output unit 211 for realizing speech synthesis of the present invention may each be implemented based on, for example, a statistical speech synthesis technique based on deep learning.
[0047] The speech learning unit 301 may include a learning text analysis unit 303 , a learning acoustic feature extraction unit 304 , and a model learning unit 305 .
[0048] The voice learning unit 301 uses, for example, data obtained by recording a singer singing a plurality of songs of an appropriate genre as learning singing voice data 312. Lyric texts of the respective songs are also prepared as learning singing voice data 311.
[0049] The learning text analysis unit 303 receives as input the learning singing data 311 including the lyrics text and analyzes the data. As a result, the learning text analysis unit 303 estimates and outputs a learning language feature value sequence 313 as a discrete numerical sequence representing the phonemes, pitches, etc. corresponding to the learning singing data 311.
[0050] The learning acoustic feature extraction unit 304 receives and analyzes learning singing voice data 312 recorded via a microphone or the like, by a singer singing lyrics corresponding to the learning singing voice data 311, in response to the input of the learning singing voice data 311. As a result, the learning acoustic feature extraction unit 304 extracts and outputs a learning acoustic feature sequence 314 representing the characteristics of the sound corresponding to the learning singing voice data 312.
[0051] In the present invention, the acoustic feature sequence corresponding to the learning acoustic feature sequence 314 and the later-described acoustic feature sequence 317 includes acoustic feature data (also referred to as formant information, spectral information, etc.) modeling the human vocal tract and vocal cord sound source data (also referred to as sound source information) modeling the human vocal cords. Spectral information can be, for example, mel cepstrum or line spectral pairs (LSP). Sound source information can be the fundamental frequency (F0) representing the pitch frequency of a human voice and its power value.
[0052] The model learning unit 305 uses machine learning to estimate the acoustic model that maximizes the probability of generating the learning acoustic feature sequence 314 from the learning language feature sequence 313. Specifically, the relationship between the text language feature sequence and the sound acoustic feature sequence is represented by a statistical model called an acoustic model. The model learning unit 305 outputs the model parameters representing the acoustic model calculated through machine learning as a learning result 315. Therefore, this acoustic model corresponds to a learned model.
[0053] As the acoustic model represented by the learning result 315 (model parameter), an HMM (Hidden Markov Model) can be used.
[0054] When a singer sings lyrics following a melody, the temporal changes in the vocal cord vibrations and characteristic parameters of the vocal tract characteristics can be learned using an HMM acoustic model. More specifically, the HMM acoustic model can be modeled using the spectrum, fundamental frequency, and their temporal structure, obtained from the learning vocal data, in phoneme units.
[0055] First, the HMM acoustic model is used Figure 3 The processing of the speech learning unit 301 is described below. The model learning unit 305 in the speech learning unit 301 can learn the HMM acoustic model with the maximum likelihood by inputting the learning language feature sequence 313 output by the learning text analysis unit 303 and the learning acoustic feature sequence 314 output by the learning acoustic feature extraction unit 304.
[0056] The spectral parameters of singing voices can be modeled using a continuous HMM. On the other hand, since the logarithmic fundamental frequency (F0) is a time series signal of variable dimensions that takes continuous values in the voiced interval and has no value in the silent interval, it cannot be directly modeled using a conventional continuous HMM or discrete HMM. Therefore, an HMM based on a probability distribution in multiple spaces corresponding to the variable dimensions, namely the MSD-HMM (Multi-Space Probability Distribution HMM), is used as the spectral parameters. The Mel-frequency cepstrum is modeled as a multidimensional Gaussian distribution, the voiced sound of the logarithmic fundamental frequency (F0) is modeled as a one-dimensional space, and the unvoiced sound is modeled as a zero-dimensional space Gaussian distribution.
[0057] Furthermore, it's known that even phonemes with identical acoustic characteristics can vary due to various factors. For example, the spectrum and logarithmic fundamental frequency (F0) of a phoneme, the fundamental unit of phonology, can vary depending on singing style and tempo, as well as the preceding and following lyrics and pitch. These factors that influence acoustic features are called context.
[0058] In one embodiment of the statistical speech synthesis process, an HMM acoustic model (context dependency model) that takes context into account can be used to accurately model the acoustic features of speech. Specifically, the learning text analysis unit 303 can output a learning language feature sequence 313 that takes into account not only the phonemes and pitch of each frame, but also the immediately preceding and following phonemes, the current position, the immediately preceding and following vibrato, and accents. Furthermore, to improve the efficiency of context combination, context clustering based on a decision tree can be used.
[0059] For example, the model learning unit 305 can generate a state continuation length decision tree as a learning result 315 for determining the state continuation length based on the learning language feature sequence 313 corresponding to the front-to-back relationship of multiple phonemes related to the state continuation length extracted by the learning text analysis unit 303 from the learning singing data 311.
[0060] In addition, the model learning unit 305 can also generate a mel-cepstral parameter decision tree for determining mel-cepstral parameters as a learning result 315 based on the learning acoustic feature sequence 314 corresponding to multiple phonemes related to the mel-cepstral parameters extracted by the learning acoustic feature extraction unit 304 from the learning singing sound data 312.
[0061] Furthermore, the model learning unit 305 may generate a logarithmic fundamental frequency decision tree for determining the logarithmic fundamental frequency (F0) as a learning result 315, based on a learning acoustic feature sequence 314 corresponding to a plurality of phonemes associated with the logarithmic fundamental frequency (F0) extracted from the learning singing voice data 312 by the learning acoustic feature extraction unit 304. Alternatively, the logarithmic fundamental frequency decision tree may be generated by modeling the voiced and unvoiced intervals of the logarithmic fundamental frequency (F0) as one-dimensional and zero-dimensional Gaussian distributions, respectively, using an MSD-HMM corresponding to a variable dimension.
[0062] Alternatively, an acoustic model based on a deep neural network (DNN) may be used in place of or in addition to an HMM-based acoustic model. In this case, the model learning unit 305 may generate model parameters representing the nonlinear transformation function of each neuron within the DNN from linguistic features to acoustic features as learning results 315. A DNN allows the relationship between a sequence of linguistic features and a sequence of acoustic features to be represented using complex nonlinear transformation functions that are difficult to represent using decision trees.
[0063] In addition, the acoustic model of the present invention is not limited to these. For example, it can also be an acoustic model that combines HMM and DNN. As long as it uses the technology of statistical sound synthesis processing, any sound synthesis method can be used.
[0064] The learning results 315 (model parameters) can be, for example, Figure 3 As shown, in Figure 1 The electronic musical instrument 10 is stored in the Figure 2 In the ROM 202 of the control system of the electronic musical instrument 10, when the power of the electronic musical instrument 10 is turned on, Figure 2 The ROM 202 is loaded into the later-described singing voice control unit 307 in the waveform data output unit 211, etc.
[0065] The learning result 315 can be, for example, Figure 3 As shown, the player operates the switch panel 140 b of the electronic musical instrument 10 , and the data is downloaded from the outside such as the Internet via the network interface 219 to the singing voice control unit 307 in the waveform data output unit 211 .
[0066] <Sound Synthesis Based on Acoustic Model>
[0067] Figure 4 This is a diagram showing an example of the waveform data output unit 211 according to one embodiment.
[0068] The waveform data output unit 211 includes a processing unit (also called a text processing unit, a pre-processing unit, etc.) 306, a singing control unit (also called an acoustic model unit) 307, a sound source 308, a singing synthesis unit (also called a pronunciation model unit) 309, etc.
[0069] The waveform data output unit 211 receives the waveform data based on the Figure 1 The keys of the keyboard 140k (performance operation member) and via Figure 2 The key scanner 206 synthesizes and outputs singing waveform data 217 corresponding to the lyrics and pitch from the singing voice data 215 containing the lyrics and pitch information and the lyrics control data instructed by the CPU 201. In other words, the waveform data output unit 211 performs the following statistical sound synthesis processing: using a statistical model such as the acoustic model set in the singing voice control unit 307 to perform prediction, the singing waveform data 217 corresponding to the singing voice data 215 containing the lyrics text is synthesized.
[0070] Furthermore, when the song data is reproduced, the waveform data output unit 211 outputs song waveform data 218 corresponding to the corresponding song reproduction position. Here, the song data may correspond to accompaniment data (e.g., data on the pitch, timbre, and sound timing of one or more notes), accompaniment and melody data, and may also be referred to as backtrack data.
[0071] The processing unit 306 receives, as a result of a player's performance (operation), a signal including the signal received by the player. Figure 2 The CPU 201 receives and parses the singing voice data 215 containing information related to the phonemes and pitch of the designated lyrics. The singing voice data 215 may include, for example, at least one of data (e.g., pitch data, note length data) of the nth note (also referred to as the nth note, nth timing, etc.), data of the nth lyric (or syllable) corresponding to the nth note, and data of the nth syllable.
[0072] For example, the processing unit 306 may determine the presence or absence of lyrics progression based on the lyric progression control method described later, based on note on / off data and pedal on / off data obtained from the operation of the keyboard 140k and pedal 140p, and obtain the singing voice data 215 corresponding to the syllable (lyrics) to be output. Furthermore, the processing unit 306 may analyze the language feature sequence 316 representing phonemes, parts of speech, words, etc. corresponding to the pitch data specified by the key or the pitch data of the obtained singing voice data 215 and the character data of the obtained singing voice data 215, and output the result to the singing voice control unit 307.
[0073] The singing data 215 may include information including at least one of the following: lyrics (characters), syllable type (starting syllable, middle syllable, ending syllable, etc.), corresponding pitch (correct pitch), and lyrics (character string) for each syllable. The singing data 215 may also include information on singing data for the nth syllable corresponding to the nth syllable (n = 1, 2, 3, 4, ...).
[0074] The vocal data 215 may also include information (such as data in a specific sound file format or MIDI data) for playing the accompaniment (song data) corresponding to the lyrics. If the vocal data is represented in the SMF format, the vocal data 215 may include a track chunk storing data related to the vocals and a track chunk storing data related to the accompaniment. The vocal data 215 may also be read from the ROM 202 into the RAM 203. The vocal data 215 is stored in a memory (e.g., the ROM 202 or RAM 203) before being played.
[0075] Lyrics control data can be as follows Figure 12 As described later, this is used to set the singing voice reproduction information corresponding to the syllable. The waveform data output unit 211 can control the timing of pronunciation based on the singing voice reproduction information. For example, the processing unit 306 can adjust the language feature value sequence 316 output to the singing voice control unit 307 based on the syllable start frame indicated by the singing voice reproduction information (for example, it can also not output frames before the syllable start frame).
[0076] The singing voice control unit 307 estimates the corresponding acoustic feature sequence 317 based on the language feature sequence 316 input from the processing unit 306 and the acoustic model set as the learning result 315, and outputs the resonance peak information 318 corresponding to the estimated acoustic feature sequence 317 to the singing voice synthesis unit 309.
[0077] For example, when the HMM acoustic model is used, the singing control unit 307 connects the HMMs according to each contextual relationship obtained from the language feature sequence 316, and refers to the decision tree, and predicts the acoustic feature sequence 317 (resonance peak information 318 and vocal sound source data 319) with the highest output probability based on the connected HMMs.
[0078] When a DNN acoustic model is used, the singing voice control unit 307 may output an acoustic feature sequence 317 in units of frames for the phoneme sequence of the language feature sequence 316 input in units of frames. Frames in the present invention may be, for example, 5 ms or 10 ms.
[0079] exist Figure 4In the process, the processing unit 306 obtains musical instrument sound data (pitch information) corresponding to the pitch of the pressed note from the memory (which may be the ROM 202 or the RAM 203 ) and outputs it to the sound source 308 .
[0080] Based on the note-on / off data input from the processing unit 306, the sound source 308 generates a sound source signal (also referred to as instrument sound waveform data) containing instrument sound data (pitch information) corresponding to the sound to be produced (note-on), and outputs it to the vocal synthesis unit 309. The sound source 308 may also perform control processing such as envelope control of the produced sound.
[0081] The vocal synthesis unit 309 forms a digital filter that models the vocal tract based on the sequence of formant information 318 sequentially input from the vocal control unit 307. Furthermore, the vocal synthesis unit 309 uses the sound source signal input from the sound source 308 as an excitation source signal, applies the digital filter, and generates and outputs the vocal waveform data 217 as a digital signal. In this case, the vocal synthesis unit 309 can also be referred to as a synthesis filter unit.
[0082] Furthermore, the singing voice synthesis unit 309 can adopt various voice synthesis methods such as the cepstrum voice synthesis method and the LSP voice synthesis method.
[0083] exist Figure 4 In the example, the output singing waveform data 217 uses the instrument sound as the sound source signal, so although it loses a little fidelity compared to the singer's singing voice, it is a singing voice that well preserves both the atmosphere of the instrument sound and the sound quality of the singer's singing voice, and effective singing waveform data 217 can be output.
[0084] Furthermore, the sound source 308 may operate so as to process the instrumental sound waveform data and output the output of other channels as the song waveform data 218. This enables an operation such as producing accompaniment sounds as normal instrumental sounds or producing a melody line instrumental sound while producing the melody of the song.
[0085] Figure 5 This is a diagram showing another example of the waveform data output unit 211 according to one embodiment. Figure 4 Repeated content will not be explained again.
[0086] Figure 5As described above, the singing voice control unit 307 estimates the acoustic feature sequence 317 based on the acoustic model. Furthermore, the singing voice control unit 307 outputs formant information 318 corresponding to the estimated acoustic feature sequence 317 and vocal cord sound source data (pitch information) 319 corresponding to the estimated acoustic feature sequence 317 to the singing voice synthesis unit 309. The singing voice control unit 307 can estimate the estimated value of the acoustic feature sequence 317 so as to maximize the probability of generating the acoustic feature sequence 317.
[0087] For example, the singing voice synthesis unit 309 may generate data (for example, singing voice waveform data of the nth lyric corresponding to the nth note) for outputting to the sound source 308 a signal obtained by applying a digital filter that models the vocal tract to a pulse train (in the case of a voiced sound phoneme) that is periodically repeated with a fundamental frequency (F0) and a power value contained in the vocal cord sound source data 319 input from the singing voice control unit 307, or white noise (in the case of an unvoiced sound phoneme) with a power value contained in the vocal cord sound source data 319, or a mixture of these signals.
[0088] The sound source 308 generates and outputs the singing voice waveform data 217 as a digital signal from the singing voice waveform data of the nth lyric corresponding to the sound to be produced (note-on), based on the note-on / off data input from the processing unit 306 .
[0089] exist Figure 5 In the example, the output singing waveform data 217 is a signal that is completely modeled by the singing control unit 307 because the sound source signal is the sound generated by the sound source 308 based on the vocal cord sound source data 319. It can output singing waveform data 217 that is very faithful to the singer's singing voice and natural.
[0090] In this way, the sound synthesis of the present invention is different from the existing sound synthesizer (vocoder) (a method of synthesizing words spoken by a person by inputting them through a microphone and replacing them with musical instrument sounds). Even if the user (performer) does not actually sing (in other words, even if the sound signal emitted by the user in real time is not input to the electronic musical instrument 10), the synthesized sound can be output through keyboard operation.
[0091] As described above, the use of statistical sound synthesis processing as a sound synthesis method allows for significantly smaller memory capacity compared to conventional segment synthesis methods. For example, electronic musical instruments using segment synthesis methods require a memory capacity of several hundred megabytes for the sound segment data. However, in this embodiment, only a few megabytes of memory are required to store the model parameters of learning result 315. This allows for lower-priced electronic musical instruments, making high-quality vocal performance systems accessible to a wider range of users.
[0092] Furthermore, in the conventional segment data method, manual adjustment of the segment data was required, resulting in a significant amount of time (on the order of years) and labor required to create data for vocal performances. However, in the present embodiment, the model parameters for the learning results 315 of the HMM acoustic model or DNN acoustic model are created with little to no data adjustment, requiring only a fraction of the time and labor. This also enables the realization of lower-priced electronic musical instruments.
[0093] Furthermore, general users can utilize the learning functions built into the server computer 300, which can be used as a cloud service, the voice synthesis LSI 205, etc. to learn their own voice, the voice of a loved one, or the voice of a celebrity, and use this as a model voice to perform singing on an electronic musical instrument. In this case, a significantly more natural and high-quality singing performance can be achieved on a more affordable electronic musical instrument than ever before.
[0094] (Lyrics Progression Control Method)
[0095] The following describes a method for controlling the lyrics progression according to an embodiment of the present invention. The lyrics progression control of the present invention may also be referred to as performance control, performance, or the like.
[0096] The main operating element (electronic musical instrument 10) in each of the following flowcharts may be referred to as the CPU 201, the waveform data output unit 211 (or its internal sound source LSI 204, the sound synthesis LSI 205 (processing unit 306, vocal control unit 307, sound source 308, vocal synthesis unit 309, etc.)), or a combination thereof. For example, the CPU 201 may execute a control processing program loaded from the ROM 202 to the RAM 203 to perform each operation.
[0097] In addition, at the beginning of the flow shown below, initialization processing can be performed. This initialization processing can include interrupt processing, exporting tick times that serve as reference times for lyrics progression and automatic accompaniment, tempo setting, song selection, song loading, instrument selection, and other processing related to buttons, etc.
[0098] The CPU 201 can detect operations of the switch panel 140 b , keyboard 140 k , pedal 140 p , etc. based on an interrupt from the key scanner 206 at appropriate timing and perform corresponding processing.
[0099] The following example shows controlling the progression of lyrics, but the objects of progression control are not limited to this. Based on the present invention, for example, the progression of arbitrary character strings, articles (e.g., press releases), etc. can also be controlled instead of lyrics. In other words, the lyrics of the present invention can also be referred to as characters, character strings, etc.
[0100] First, the present invention provides an overview of the method for controlling the syllable positions of lyrics (also referred to as phrases, etc.). This control method allows for quick and intuitive control of lyrics using a keyboard. Furthermore, the present invention assumes that a "syllable" represents a single word (or character) such as "go," "for," or "it," and that "lyrics" or "phrases" represent a phrase (or article) consisting of multiple syllables or multiple words (or multiple characters) such as "Go for it." However, these definitions may differ.
[0101] Furthermore, in the present invention, a syllable position can be represented by a specific index (e.g., a syllable index). The syllable index can be a variable indicating the syllable (or character) corresponding to the syllable number (or character number) from the beginning of the syllables in the lyrics. In the present invention, the syllable position and the syllable index can be interchangeably referred to.
[0102] In the present invention, the lyrics corresponding to one syllable index may correspond to one or more characters constituting one syllable. A syllable may include various syllables such as only vowels, only consonants, or both consonants and vowels.
[0103] Figure 6 This figure shows an example of a keyboard range division for syllable position control according to one embodiment. In this example, the keyboard 140k is divided into a first key range (first range) and a second key range (second range). That is, the keyboard 140k has a plurality of keys including a plurality of first keys corresponding to the first range and a plurality of second keys corresponding to the second range. In addition, in this example, an example in which the number of keys of the keyboard 140k is 61 is shown, but the embodiments of the present invention can also be similarly applied to cases with other numbers of keys.
[0104] In the present invention, the key range may be interchangeably referred to as a region (or range) of a keyboard, a region (or range) of a performance operating element, a musical range, or a region (or range) of a sound.
[0105] The first key range can also be called the syllable position control key range, keyboard control key range, control key range, etc., and is used to specify the syllable position. In other words, the control key range may not be used to specify the pitch, strength, length, etc. of the played note.
[0106] For example, the control key range can correspond to the key range for chord articulation (e.g., C1-F2). Within the control key range, the keys used to control syllable positions can be composed solely of white keys, solely of black keys, or a combination of both. For example, if only white keys are used to control syllable positions, the black keys within the control key range can be used to control lyrics (e.g., transitioning to a later / previous lyric in a particular song).
[0107] The second key range, which may also be referred to as the keyboard playing key range or the playing key range, is used to specify pitch, velocity, duration, etc. The electronic musical instrument 10 uses the pitch (interval), velocity, etc. specified by operating the playing key range to produce a sound corresponding to the syllable position (or lyrics) specified by operating the control key range.
[0108] In addition, Figure 6 In the example shown, the control range consists of several keys on the left hand side, and the performance range consists of keys that do not correspond to the control range. However, the present invention is not limited to this. For example, each key range may be composed of non-adjacent (dispersed) keys, or the control range may be composed of keys on the right hand side, and the performance range may be composed of keys on the left hand side.
[0109] Figures 7A to 7C This is a diagram showing an example of syllables assigned to the control key range. Figure 7A This shows an example of lyrics for which syllable positions are controlled using the control key range. The lyrics for "まばたきしてはみんなを" are shown. The pitch and length of a note are used as examples, and the actual sound output can be controlled using the performance key range.
[0110] Figure 7B Indicates that Figure 7A An example of assigning each syllable of the lyrics to the white keys in the control key range. In this example, a total of 11 white keys, C1 to F2 in the control key range, are mapped to each of the 1 syllables of the lyrics.
[0111] When a white key in the control key range is pressed, the electronic musical instrument 10 sets the syllable position to the position corresponding to the white key (for example, if the white key is G1, it is set to "し"). When C1 is pressed, the electronic musical instrument 10 starts the lyrics at the beginning (sets the syllable position to "ま") regardless of the current syllable position.
[0112] The electronic musical instrument 10 shifts the syllable position by one (moves to the next) when any key in the performance key range is pressed while no key in the control key range is pressed (for example, if the position before the key is pressed is "ま", it shifts to "ば"). In addition, when the syllable position reaches the end of the lyrics, the syllable position can be changed to the beginning of the lyrics ( Figure 7B It can also be changed to the beginning of the next lyric of the lyric.
[0113] The electronic musical instrument 10, when a white key in the control key range is pressed, maintains the syllable position corresponding to the white key unchanged even if any key in the performance key range is pressed multiple times (for example, if the position corresponding to the white key is "し", "し" is pronounced every time a key in the performance key range is pressed).
[0114] When a white key in the control range is pressed, the electronic musical instrument 10 can also pronounce the syllable corresponding to the white key based on the key pressed in the performance range. For example, if a key in the performance range is pressed in the order C2 → D1 → E1 in the control range, the electronic musical instrument 10 can pronounce "みばた" at the pitch corresponding to the key in the performance range. This operation allows the syllables of the lyrics corresponding to the control range to be pronounced in any order (freely creating anagrams).
[0115] Figure 7C This example shows how to assign syllables from other lyrics (English lyrics) to the white keys within the control key range. In this example, the 11 white keys (C1-F2) within the control key range are mapped to the syllables of the lyrics "holy infant so tender and mild sleep in." This allows you to assign syllables in any language.
[0116] For 1 key, you can either Figure 7B 、 7C As shown, one character is assigned to one syllable, but multiple characters and multiple syllables may be assigned.
[0117] Lyrics and syllable related data may also correspond to the singing voice data 215 (also referred to as lyrics data, syllable data, etc.). For example, the electronic musical instrument 10 may store multiple lyrics data in the memory, and select one lyric data when a specific function key (e.g., button, switch, etc.) is operated.
[0118] <Lyrics Progress Control>
[0119] Figure 8This is a diagram showing an example of a flowchart of a lyrics progression control method according to one embodiment.
[0120] First, the electronic musical instrument 10 sets the note position control flag to "invalid" as an initial value (step S101).
[0121] The electronic musical instrument 10 determines whether syllable allocation is required (step S102). The electronic musical instrument 10 may determine that syllable allocation is required when a specific function key (e.g., button, switch, etc.) of the electronic musical instrument 10 is operated (and lyrics are loaded, etc.).
[0122] If syllable allocation is required ("Yes" in step S102), the electronic musical instrument 10 performs syllable allocation processing on the control key range (white keys) (step S103) and sets the syllable position control flag to "valid" (step S104). The syllable to be allocated can be selected from a plurality of lyrics data as described above. The syllable position control flag "valid" can also be referred to as "keyboard division valid."
[0123] If syllable allocation is not required ("No" in step S102), the control key range is not set and all keys are used for pitch specification (normal performance mode). The syllable position control flag "invalid" can also be referred to as keyboard division invalid.
[0124] After step S104 or step S102 returns "No", the electronic musical instrument 10 determines whether any keyboard operation has occurred (step S105). If a keyboard operation has occurred (step S105 returns "Yes"), the electronic musical instrument 10 obtains information (also referred to as key-press / key-release information) such as the key that has been pressed / is being pressed and the key that has been released / is being released (step S106).
[0125] After step S106, the electronic musical instrument 10 checks whether the syllable position control flag is valid (step S107). If the syllable position control flag is valid ("Yes" in step S107), the syllable position control process is performed (step S108). Otherwise ("No" in step S107), the electronic musical instrument 10 performs the performance control process (step S109). Figure 9 The performance control process is described later in Figure 10 To be discussed later.
[0126] After step S108 or step S109, the electronic musical instrument 10 determines whether the playback of the lyrics has ended (step S110). If so ("Yes" in step S110), the electronic musical instrument 10 may terminate the process in this flowchart and return to standby mode. Otherwise ("No" in step S110), the process may return to step S102 or step S105. The term "whether the playback of the lyrics has ended" here may refer to the playback of the lyrics of a single phrase or the entire song.
[0127] <Syllable Position Control>
[0128] Figure 9 This is a diagram showing an example of a flowchart of a syllable position control process according to one embodiment.
[0129] The electronic musical instrument 10 determines whether a key press / release operation is performed in the control key range (step S201). If an operation is performed in the control key range ("Yes" in step S201), it determines whether the operation is a key press (step S202).
[0130] If a key operation is performed ("Yes" in step S202), the electronic musical instrument 10 saves (or stores or sets) the information of the key pressed by the key operation as a syllable control key (step S203). Furthermore, the electronic musical instrument 10 resets (or does not set) the key release flag (step S204). The key release flag is reset if any key in the control key range is pressed, and is set otherwise.
[0131] On the other hand, when a key-release operation is performed (No in step S202), the electronic musical instrument 10 determines whether the information of the key released by the key-release operation is the same as the stored syllable control key (step S205).
[0132] If the information of the released key is identical to the stored syllable control key ("Yes" in step S205), the key release flag is set (step S206). Alternatively, if the information of the released key is identical to the stored syllable control key, but there is a key in the control key range, the electronic musical instrument 10 may store the information of the key in the range as a syllable control key, in which case the key release flag may not be set.
[0133] On the other hand, if there is no operation in the control key range (No in step S201), the electronic musical instrument 10 performs a performance control process (step S207). The performance control process in step S207 may be the same as the performance control process in step S109.
[0134] After step S204 , step S206 , “No” in step S205 , or step S207 , the electronic musical instrument 10 may end the syllable position control process.
[0135] The syllable control key, also referred to as syllable control information, may be information about the key number of a pressed / released key or information about the pitch (or note number) of a pressed / released key. The present invention will be described below using the example of a syllable control key holding a key number, but the present invention is not limited thereto.
[0136] In addition, for example, Figure 7B and Figure 7C In the example, the keys corresponding to C1 to F2 may correspond to key numbers 0 to 11, respectively. The key number may also be a string representing a pitch (eg, C1, F2).
[0137] according to Figure 9 In the syllable position control process, if a key is pressed in the control key range, the key is held. If a key is released in the control key range, a key release flag is set while the held key is maintained. When another key in the control key range is pressed, the held key is replaced by the other key. Alternatively, if a new key is pressed while no key in the control key range has been released, the held key may be overwritten by the new key.
[0138] <Performance Control>
[0139] Figure 10 This is a diagram showing an example of a flowchart of a performance control process according to one embodiment.
[0140] The electronic musical instrument 10 performs a syllable progression determination process (step S301). The syllable progression determination process returns a determination result (return value) as to whether the syllable position is to be advanced. If the determination result is "yes" (or "true"), the current syllable position is obtained, and the syllable position is shifted (or displaced, advanced) by one (in other words, the lyrics are advanced) (step S302). An example of a syllable progression determination process is shown in FIG. Figure 11 To be discussed later.
[0141] On the other hand, if the determination result of the syllable progression determination process in step S301 is "No" (or "False"), the syllable position is not changed.
[0142] After step S302, the electronic musical instrument 10 determines whether the syllable control key is set (holding a valid value) (step S303). If the syllable control key is set ("Yes" in step S303), the electronic musical instrument 10 determines whether the syllable control key is a syllable position designation valid key (also referred to as a valid key) (step S304).
[0143] Here, valid keys may refer to keys assigned syllables among all keys in the control key range. For example, if the number of syllables in the current lyrics is less than the number of white keys in the control key range, some of the white keys in the control key range will correspond to valid keys, while the remaining ones will not. Furthermore, in this case, the black keys will not correspond to valid keys either.
[0144] As can be seen from this, if the lyrics change, which key becomes the effective key and also can change. In addition, it is not necessary for 1 key to correspond to 1 syllable in a one-to-one manner, but also 1 key can correspond to multiple syllables, or multiple keys can correspond to 1 syllable.
[0145] If the syllable control key is a valid key ("YES" in step S304), the electronic musical instrument 10 obtains the syllable position corresponding to (the key number of) the syllable control key (step S305).
[0146] After step S305, the electronic musical instrument 10 determines whether the key release flag is set (step S306). If the key release flag is set ("Yes" in step S306), the electronic musical instrument 10 clears the syllable control key (or sets an invalid value) (step S307).
[0147] After "No" in step S303, "No" in step S304, "No" in step S306, or step S307, the electronic musical instrument 10 performs a syllable change process (step S308). Figure 12 In addition, as will be described later, a syllable performance (reproduction) process can be performed during the syllable change process.
[0148] Alternatively, before or after the syllable change process, the electronic musical instrument 10 may store the current syllable position (the syllable position obtained in step S302 or step S305 (or the syllable position obtained and advanced by one)) as the current syllable position in the storage unit. Acquisition of the syllable position in step S302 may be acquisition of the stored current syllable position. Furthermore, instead of advancing the syllable position by one in step S302, the syllable position may be advanced by one before or after the syllable change process in step S308.
[0149] In response to “No” in step S301 or after step S308 , the electronic musical instrument 10 may end the performance control process.
[0150] <Syllable Progression Discrimination>
[0151] Figure 11This figure shows an example of a flowchart for syllable progression discrimination processing according to one embodiment. In other words, this processing corresponds to the following: if a single note is pressed in the performance key range, the syllable is progressed; further, if a chord is pressed in the performance key range, the syllable progression is determined based on which pitch (or, alternatively, "which pitch," "which part," etc.) of the chord is changed by the key press.
[0152] The electronic musical instrument 10 obtains the current number of keys pressed in the performance key range (step S401 ).
[0153] Next, the electronic musical instrument 10 determines whether the current number of keys in the performance key range is 2 or more (whether there are 2 or more keys) (step S402). If the current number of keys is 2 or more ("Yes" in step S402), the electronic musical instrument 10 obtains the key press time and key number corresponding to each key (step S403).
[0154] After step S403, the electronic musical instrument 10 determines whether the difference between the most recent key press time and the last key press time in the performance key range is within the harmony discrimination time (step S404). Step S404 can also be considered a step of determining whether the difference between the key press time of the newly pressed note and the last key press time (or the key press time i times ago (i is an integer)) is within the harmony discrimination time. Preferably, the past key press time corresponds to a key that was continuously pressed during the most recent key press time.
[0155] Here, the harmony determination time is the time (period) used to determine that multiple notes emitted within that time are simultaneous harmonies, and that multiple notes emitted outside that time are independent notes (e.g., notes of a melody line) or discrete harmonies. The harmony determination time can be expressed, for example, in milliseconds or microseconds.
[0156] The harmony identification time may be obtained by user input or derived based on the tempo of the music. The harmony identification time may also be referred to as a predetermined set time, a set time, or the like.
[0157] If the difference between the latest key press time and the previous key press time is within the harmony determination time ("YES" in step S404), the electronic musical instrument 10 determines that the key pressed note is a simultaneous harmony (harmony is specified). Furthermore, the electronic musical instrument 10 determines that the syllable is maintained (the lyrics are not progressed), and sets the return value of the syllable progression determination process to "NO" (or "FALSE") (step S405).
[0158] According to the determination in step S404 , when a plurality of keys are pressed with the intention of harmony, it is possible to advance only one syllable in response to the fact that it is not desired to advance the syllable by the number of keys.
[0159] On the other hand, if the key-pressing time has not elapsed within the harmony discrimination time ("No" in step S404), it is determined whether the current number of key-pressings in the performance key range is greater than or equal to a predetermined number and whether the most recently pressed tone (key) corresponds to a specific tone (key) among all the keys pressed in the performance key range (step S406). Furthermore, if the answer is "No" in step S404, the electronic musical instrument 10 may determine that the harmony designation has been released or that no harmony designation has been made.
[0160] The predetermined number may be, for example, 2, 4, or 8. Furthermore, the specific note (key) may be the lowest note (key) among all the notes (keys) pressed, or the ith (i is an integer) highest or lowest note (key). These predetermined numbers and specific notes may be set by user operation or may be predetermined in advance.
[0161] If the answer is “YES” in step S406 , the electronic musical instrument 10 determines that the syllable is to be advanced (the lyrics are to be advanced), and sets the return value of the syllable advancement determination process to “YES” (or “TRUE”) (step S407 ).
[0162] If step S406 is "No", the electronic musical instrument 10 determines that although it is not simultaneous harmony, the syllable is maintained (the lyrics are not progressed), and sets the return value of the syllable progress determination process to "No" (or "False") (step S405).
[0163] If step S402 is "No", the electronic musical instrument 10 determines that the syllable is advanced (the lyrics are advanced) because it is not simultaneous harmony, and sets the return value of the syllable advancement determination process to "Yes" (or "True") (step S407).
[0164] according to Figure 11 Such syllable progression determination processing can make the syllable progress if, for example, multiple tones are pronounced with a large time difference (melody) rather than multiple tones pronounced with a small time difference (so-called harmony), the syllable progresses.
[0165] <Syllable Change>
[0166] Figure 12 This is a diagram showing an example of a flowchart of a syllable change process according to one embodiment.
[0167] The electronic musical instrument 10 obtains lyrics control data corresponding to the syllable positions obtained in the performance control process (step S501 ).
[0168] Here, the lyrics control data may be data containing parameters related to the pronunciation (singing synthesis) of each syllable contained in the lyrics. If data containing parameters related to the pronunciation of a syllable is called syllable control data, the lyrics control data may include more than one syllable control data.
[0169] For example, syllable control data can include information such as pronunciation timing, syllable start frame, vowel start frame, vowel end frame, syllable end frame, and (character information of) lyrics (or syllables). Furthermore, a frame can be the unit of construction of the aforementioned phonemes (phoneme sequences), or it can be replaced by another time unit. Hereinafter, lyrics control data and syllable control data will not be specifically distinguished.
[0170] The pronunciation timing can represent the timing (or offset) that serves as the basis for each frame (e.g., syllable start frame, vowel start frame, etc.). This pronunciation timing can be expressed as the time from the key press. The pronunciation timing and each frame information can also be specified in frame numbers (frame units).
[0171] The sound corresponding to a syllable may start to be pronounced at the syllable start frame and end to be pronounced at the syllable end frame. The sound corresponding to a vowel in a syllable may start to be pronounced at the vowel start frame and end to be pronounced at the vowel end frame. That is, generally, the vowel start frame has a value greater than the syllable start frame, and the vowel end frame has a value less than the syllable end frame.
[0172] The syllable start frame may correspond to the beginning address information of the frame of the syllable, and the syllable end frame may correspond to the end address information of the frame of the syllable.
[0173] Next, the electronic musical instrument 10 determines whether the syllable start frame of the lyrics control data obtained in step S501 needs to be adjusted (step S502). For example, the electronic musical instrument 10 may determine that the syllable start frame needs to be adjusted if a frame position adjustment flag is set. The electronic musical instrument 10 may control the value of the frame position adjustment flag based on the operation of a function key or may determine the value of the frame position adjustment flag based on parameters in the lyrics control data.
[0174] If the syllable-onset frame needs to be adjusted ("Yes" in step S502), the electronic musical instrument 10 adjusts the syllable-onset frame based on the adjustment coefficient (step S503). For example, the electronic musical instrument 10 may calculate a value obtained by applying a predetermined operation (e.g., addition, subtraction, multiplication, or division) using the adjustment coefficient to the syllable-onset frame as the new (adjusted) syllable-onset frame.
[0175] The adjustment coefficient may be a parameter (e.g., offset, frame number, etc.) suitable for reducing (or removing) the white noise portion of a syllable. The adjustment coefficient may have a different (or independent) value for each syllable. The adjustment coefficient may be included in the lyrics control data or determined based on the lyrics control data.
[0176] Furthermore, the adjustment of the syllable start frame in step S503 may be applied only to the tones pronounced when a key is pressed in the control key range, or may be applied to the tones pronounced when no key is pressed in the control key range.
[0177] After step S503, the electronic musical instrument 10 determines whether the adjusted syllable onset frame value is greater than the vowel onset frame value (step S504). If the adjusted syllable onset frame value is greater than the vowel onset frame value ("Yes" in step S504), the electronic musical instrument 10 changes the adjusted syllable onset frame value to the vowel onset frame value (step S505).
[0178] According to steps S504 and S505, for example, pronunciation can be started from the beginning of a vowel while minimizing white noise. If pronunciation is started midway through a vowel, the attack of the pronunciation will be deteriorated. However, starting pronunciation from the beginning of a vowel can suppress the deterioration of the attack.
[0179] After "No" in step S502, "No" in step S504, or step S505, the electronic musical instrument 10 sets information including at least a syllable start frame, a vowel start frame, a vowel end frame, and a syllable end frame as singing voice reproduction information (step S506). As described above, the syllable start frame here can be the value of the syllable start frame included in the lyrics control data, the value of the syllable start frame adjusted using the adjustment coefficient, or the value of the vowel start frame.
[0180] The electronic musical instrument 10 applies the singing voice reproduction process and produces the sound corresponding to the current syllable position (step S507). In the singing voice reproduction process, the electronic musical instrument 10 can produce the sound corresponding to the current syllable position based on the singing voice reproduction information of step S506 and the key pressed in the performance key range (the pitch, etc.).
[0181] In the singing reproduction process, the electronic musical instrument 10 can also obtain the acoustic feature data (resonance peak information) of the singing data corresponding to the current syllable position through the singing control unit 307, instruct the sound source 308 to pronounce the instrument sound of the pitch corresponding to the key (generation of instrument sound waveform data), and instruct the singing synthesis unit 309 to assign the above-mentioned resonance peak information to the instrument sound waveform data output from the sound source 308.
[0182] For example, the processing unit 306 inputs the designated pitch data (pitch data corresponding to the pressed key), singing voice data corresponding to the current syllable position, and singing voice reproduction information corresponding to the current syllable position to the singing voice control unit 307. The singing voice control unit 307 estimates an acoustic feature value sequence 317 based on the input and outputs the corresponding formant information 318 and vocal cord sound source data (pitch information) 319 to the singing voice synthesis unit 309. This acoustic feature value sequence 317 can adjust the reproduction start frame based on the singing voice reproduction information.
[0183] The singing voice synthesis unit 309 generates singing voice waveform data based on the input formant information 318 and vocal cord sound source data (pitch information) 319, and outputs the data to the sound source 308. The sound source 308 then performs sound production processing on the singing voice waveform data received from the singing voice synthesis unit 309.
[0184] In addition, when the judgment result of the syllable progression judgment processing in step S301 of the electronic musical instrument 10 is "No" (or "False"), the singing reproduction processing is applied based on the singing reproduction information that has been obtained and the keys pressed in the performance key range to pronounce the sound corresponding to the current syllable position.
[0185] <Modification>
[0186] Alternatively, in the electronic musical instrument 10, at least one of characters, graphics, patterns, or designs may be displayed on a key to which a syllable within a control key range is assigned, so that the assigned syllable can be visually identified (or distinguished, grasped, understood), and at least one of the color, brightness, and chromaticity of the key (for example, a light-emitting element (LED: Light Emitting Diode) built into the key) may be changed.
[0187] In addition, in the electronic musical instrument 10, at least one character, graphic, pattern, or design that is different from other keys may be displayed on the key corresponding to the current syllable position, so that the current syllable position can be visually identified (or distinguished, grasped, understood) (in other words, it can be distinguished from other keys), and at least one key color, brightness, and chromaticity that is different from other keys may also be displayed.
[0188] Figure 13A and Figure 13B This figure shows an example of the appearance of the keys of the control key range. In this example, the lyrics of "まばたきしてはみんなを" are displayed in a visually recognizable manner on each of the 11 white keys C1 to F2 in the control key range.
[0189] In addition, Figure 13A In the figure, a portion of the C1 bond emits light (the "0" portion in the figure). Figure 13BIn the figure, a part of the D1 key is illuminated (the "0" part in the figure). Figure 13A and Figure 13B , respectively, to make it easy for the player to understand that the current syllable position is "ま" or "ば".
[0190] In addition, in Figure 13A and Figure 13B When a display is provided that allows for easy understanding of the keys to which syllables are assigned, the number of keys in the control key range can be varied depending on the lyrics currently being played. For example, if the number of syllables in the lyrics is x (x is an integer), the control key range only needs to include x white keys. This prevents the situation where the number of keys in the performance key range is always small (limiting the degree of freedom in playable pitches) regardless of the selected lyrics.
[0191] In the above embodiment, it is envisioned that lyrics data is selected based on the operation of a specific function key (e.g., a button, switch, etc.), but the present invention is not limited to this. For example, the electronic musical instrument 10 may also select lyrics data based on the operation of a key (e.g., a black key) within the control key range that is not assigned a syllable. For example, the leftmost black key within the control key range may select the lyrics that are one lyric before the current lyric, and the second black key from the left within the control key range may select the lyrics that are one lyric after the current lyric.
[0192] The electronic musical instrument 10 can also control the display 150d to display lyrics. For example, the lyrics near the current lyric position (syllable index) can be displayed, or the lyrics corresponding to the currently pronounced notes or the lyrics corresponding to the pronounced notes can be displayed by coloring, etc., so that the current lyric position can be identified.
[0193] The electronic musical instrument 10 may also transmit at least one of the following: singing data, information about the current location of lyrics, etc., to an external device (e.g., a smartphone or tablet terminal). The external device may control its own display to display the lyrics based on the received singing data, information about the current location of lyrics, etc.
[0194] In the above example, the electronic musical instrument 10 is a keyboard instrument such as a keyboard, but the present invention is not limited thereto. The electronic musical instrument 10 can be any device having a structure that allows the user to specify the timing of sound production, and can also be an electric violin, electric guitar, drums, trumpet, etc.
[0195] Therefore, the term "key" in the present invention can also be changed to "string," "tube," other performance operating element for specifying pitch, or any other performance operating element. The term "key" in the present invention can also be changed to "key striking," "plucking," "performance," "operation of an operating element," or "user operation." The term "releasing a key" in the present invention can also be changed to "stopping a string," "silencing," "stopping performance," or "stopping (non-operation)" of an operating element.
[0196] Furthermore, the operating elements (e.g., playing elements, keys) of the present invention may also be operating elements (key images, etc.) displayed on a touch panel, a virtual keyboard, etc. In this case, the electronic musical instrument 10 is not limited to a so-called musical instrument (keyboard, etc.), but may also be a mobile phone, a smartphone, a tablet terminal, a personal computer (PC), a television, etc.
[0197] Figure 14 This diagram illustrates an example of a tablet terminal implementing a lyric progression control method according to one embodiment. The tablet terminal 10t displays at least a keyboard 140k on its display. A portion of the keyboard 140k (in this example, the 11 white keys C1-F2) corresponds to the control key range. The lyrics of "まばたきしてはみんなを" are displayed on each of the 11 white keys C1-F2 within the control key range for visual recognition.
[0198] Alternatively, an external device that receives the above-mentioned singing data and information related to the current position of the lyrics may also display the Figure 14 As shown, the assigned syllables, the keyboard 140k indicating the current syllable position, etc.
[0199] As described above, the electronic musical instrument 10 of the present invention can provide a new performance experience and enable the user (player) to enjoy the performance more.
[0200] For example, the electronic musical instrument 10 of the present invention can easily start the lyrics. Since the position of the syllable is visually known, it is possible to directly jump to any syllable during the performance of the lyrics.
[0201] Furthermore, the electronic musical instrument 10 of the present invention can directly designate and maintain any vowel using only the keyboard when a syllable (vowel) is desired to be held at a specific syllable position during lyric performance, and can also perform melisma performance without using pedals or buttons.
[0202] Furthermore, the electronic musical instrument 10 of the present invention can randomly change syllable positions based on keyboard operation, allowing for performance while changing syllable combinations. Therefore, not only can original lyrics be created, but also alternative lyrics, such as anagrams, can be produced. For example, when combined with automatic performance tools such as looping or an arpeggiator, a new performance experience can be provided, generating lyrics and phrases beyond the user's expectations.
[0203] In addition, the electronic musical instrument 10 can have a plurality of performance operating members (such as keys) and a processor (such as a CPU) that respectively establish corresponding pitch data different from each other. The above-mentioned processor can determine the syllable position included in the phrase based on the operation (such as pressing / releasing a key) of the performance operating member included in the first range (control key range) of the above-mentioned plurality of performance operating members. In addition, the above-mentioned processor can also indicate the pronunciation of the syllable corresponding to the above-mentioned syllable position determined based on the operation of the performance operating member included in the second range (performance key range) of the above-mentioned plurality of performance operating members. According to such a structure, for example, the part of the lyrics that the user wants to pronounce can be easily specified using only the keyboard.
[0204] Furthermore, the processor may determine the syllable position based on the key number corresponding to the operated performance operating member included in the first range when the performance operating member included in the first range is operated. With such a configuration, any syllable can be intuitively changed by pressing a key in the first range.
[0205] Furthermore, the processor may determine the syllable position based on the key number corresponding to the operated performance operating member in the first range when the performance operating member in the first range is operated and the operated performance operating member in the first range is a valid key to which a syllable is assigned. With this configuration, it is possible to intuitively change to any syllable by operating a key in the first range to which a syllable is assigned. Keys not to which a syllable is assigned can be used for purposes other than syllable change.
[0206] Furthermore, the processor may shift the syllable position by one based on an operation of a performance operating member included in the second range, when the performance operating member included in the first range is not operated. This configuration enables user-friendly operation such as advancing syllables primarily through operations in the second range and skipping syllables by operating the first range only when necessary.
[0207] Furthermore, the processor may instruct the pronunciation of the syllable start frame of the syllable corresponding to the syllable position adjusted based on the adjustment coefficient. According to such a configuration, the white noise portion of the syllable can be appropriately reduced (or removed).
[0208] Furthermore, if the value of the syllable start frame adjusted based on the adjustment coefficient is greater than the value of the vowel start frame of the syllable, the processor may set the adjusted syllable start frame value to be the same as the value of the vowel start frame. With this configuration, it is possible to minimize white noise while suppressing degradation of the sense of attack.
[0209] Furthermore, the processor may control the syllable to be pronounced without advancing, while operation of a performance operating member included in the first range among the plurality of performance operating members continues, regardless of how a performance operating member included in the second range among the plurality of performance operating members is operated, and control the syllable to be pronounced whenever a performance operating member included in the second range is operated, while no operation of a performance operating member included in the first range is performed. Furthermore, the processor may instruct the pronunciation of the syllable corresponding to the syllable position at a pitch specified based on the operation of the performance operating member included in the second range. With such a configuration, syllable maintenance can be easily achieved.
[0210] Furthermore, the processor may be configured to control the syllable to not advance from the position of the syllable corresponding to the syllable in the first range during which the operation is continued, regardless of how the syllable in the second range is operated, while the operation of the syllable in the first range continues. This configuration enables user-friendly operation, such as advancing syllables primarily through operations in the second range and skipping syllables by operating the first range only when necessary.
[0211] Furthermore, each syllable included in the phrase may be assigned to each performance operation element included in the first range. With such a configuration, the user can easily grasp the current syllable position.
[0212] Alternatively, the processor may use the performance operating elements within the first range to determine the position of the syllable when a specific function key is operated by the user, and otherwise use the performance operating elements within the first range to specify the pitch of the sound to be produced (normal mode, normal performance action). This configuration allows for appropriate control of whether or not lyrics progression control utilizing keyboard division is enabled.
[0213] Furthermore, the processor may apply a display for the user to understand the assigned syllables to the performance operating elements included in the first range. With this configuration, the user can easily understand the keys corresponding to the syllables constituting the lyrics, thereby appropriately facilitating subsequent user operations.
[0214] Furthermore, the processor may control the external device to transmit information for causing the external device to display information for the user to understand the syllables assigned to the performance operation elements included in the first musical range. With this configuration, the user can easily grasp the keys corresponding to the syllables constituting the lyrics by visually recognizing the external device, thereby facilitating subsequent user operations.
[0215] In addition, the block diagrams used in the description of the above embodiments represent blocks of functional units. These functional blocks (constituent parts) are implemented by any combination of hardware and / or software. In addition, the means for implementing each functional block are not particularly limited. That is, each functional block can be implemented by a physically combined device, or two or more physically separated devices can be connected by wired or wireless means and implemented by these multiple devices.
[0216] Furthermore, terms used in the description of the present invention and / or terms necessary for understanding the present invention may be replaced with terms having the same or similar meanings.
[0217] The information, parameters, etc. described in the present invention may be expressed in absolute values, relative values relative to specified values, or other corresponding information. In addition, the names used for parameters, etc. in the present invention are not restrictive in any way.
[0218] The information, signals, etc. described in the present invention may also be represented using any of a variety of different technologies. For example, the data, commands, instructions, information, signals, bits, symbols, chips, etc. mentioned in the above overall description may also be represented by voltage, current, electromagnetic waves, magnetic fields or magnetic particles, light fields or photons, or any combination thereof.
[0219] Information, signals, etc. can also be input and output through multiple network nodes. Input and output information, signals, etc. can be stored in a specific location (e.g., memory) or managed using a table. Input and output information, signals, etc. can be overwritten, updated, or appended. Output information, signals, etc. can also be deleted. Input information, signals, etc. can also be sent to other devices.
[0220] Software, whether referred to as software, firmware, middleware, microcode, hardware description language, or by other names, shall be construed broadly to mean instructions, sets of instructions, code, code segments, program codes, programs, subroutines, software modules, applications, software applications, software packages, routines, subroutines, objects, executables, execution threads, steps, functions, or the like.
[0221] Furthermore, software, commands, information, and the like may also be transmitted and received via a transmission medium. For example, if software is transmitted from a website, server, or other remote source using at least one of a wired technology (coaxial cable, fiber optic cable, twisted pair, digital subscriber line (DSL), etc.) and a wireless technology (infrared, microwave, etc.), at least one of these wired and wireless technologies is included in the definition of a transmission medium.
[0222] The various aspects / implementations described in the present invention may be used individually or in combination, or may be switched during execution. Furthermore, the processing steps, sequences, flow charts, and the like in the various aspects / implementations described in the present invention may be reordered as long as there are no inconsistencies. For example, the methods described in the present invention describe various step elements using illustrative sequences, and are not limited to the specific sequences described.
[0223] The phrase "based on" used in the present invention does not mean "based only on" unless otherwise specified. In other words, the phrase "based on" means both "based only on" and "based at least on."
[0224] Any reference to an element using a designation such as "first" or "second" in the present invention does not limit the quantity or order of these elements. These designations may be used in the present invention as a convenient method for distinguishing between two or more elements. Therefore, reference to a first and a second element does not mean that only two elements can be used or that the first element must precede the second element in some manner.
[0225] In the present invention, when the terms "include," "including," and their variations are used, these terms, like the term "comprising," are intended to be inclusive. Furthermore, the term "or" used in the present invention does not mean an exclusive logical AND.
[0226] "A / B" in the present invention may also mean "at least one of A and B".
[0227] In the present invention, for example, when an article is added by translation like a, an, and the in English, the present invention also includes the case where the noun following the article is in plural form.
[0228] While the invention of this disclosure has been described in detail above, it will be apparent to those skilled in the art that the invention of this disclosure is not limited to the embodiments described herein. The invention of this disclosure can be implemented in the form of modifications and variations without departing from the spirit and scope of the invention as set forth in the claims. Therefore, the description of this disclosure is for illustrative purposes only and is not intended to limit the invention of this disclosure in any way.
Claims
1. An electronic musical instrument, characterized in that have: a plurality of keys, including a plurality of first keys corresponding to the first range and a plurality of second keys corresponding to the second range; and At least 1 processor; The at least one processor determines the position of syllables included in the phrase based on the key operation of the first range; The at least one processor instructs the pronunciation of the sound corresponding to the determined syllable position according to the key operation of the second range. The at least one processor instructs pronunciation of a syllable start frame of a syllable corresponding to the syllable position after adjusting the syllable start frame based on the adjustment coefficient.
2. The electronic musical instrument according to claim 1, wherein The at least one processor determines the syllable position based on a key number corresponding to the operated key in the first range.
3. The electronic musical instrument according to claim 2, wherein The at least one processor determines the syllable position when the operated key of the first range is a valid key to which a syllable is assigned.
4. The electronic musical instrument according to any one of claims 1 to 3, wherein The at least one processor shifts the syllable position by one based on the key operation in the second range when the key in the first range is not operated.
5. The electronic musical instrument according to claim 1, wherein The at least one processor sets the adjusted syllable start frame value to the same as the vowel start frame value when the syllable start frame value adjusted based on the adjustment coefficient is larger than the vowel start frame value of the syllable.
6. A lyrics progression control method, characterized in that: At least one processor of an electronic musical instrument including a plurality of first keys corresponding to a first musical range and a plurality of second keys corresponding to a second musical range determines a syllable position included in a musical phrase based on a key operation in the first musical range, and instructs pronunciation of a note corresponding to the determined syllable position based on a key operation in the second musical range. The at least one processor instructs pronunciation of a syllable start frame of a syllable corresponding to the syllable position after adjusting the syllable start frame based on the adjustment coefficient.
7. The lyrics progression control method according to claim 6, wherein: The at least one processor determines the syllable position based on a key number corresponding to the operated key in the first range.
8. The lyrics progression control method according to claim 7, wherein: The at least one processor determines the syllable position when the operated key of the first range is a valid key to which a syllable is assigned.
9. The lyrics progression control method according to any one of claims 6 to 8, characterized in that: The at least one processor shifts the syllable position by one based on the key operation in the second range when the key in the first range is not operated.
10. The lyrics progression control method according to claim 6, wherein: The at least one processor sets the adjusted syllable start frame value to the same as the vowel start frame value when the syllable start frame value adjusted based on the adjustment coefficient is larger than the vowel start frame value of the syllable.
11. A recording medium, characterized in that Contains the following programs: At least one processor of an electronic musical instrument including a plurality of first keys corresponding to a first musical range and a plurality of second keys corresponding to a second musical range determines a syllable position included in a musical phrase based on a key operation in the first musical range, and instructs pronunciation of a note corresponding to the determined syllable position based on a key operation in the second musical range. The at least one processor instructs pronunciation of a syllable start frame of a syllable corresponding to the syllable position after adjusting the syllable start frame based on the adjustment coefficient.
12. A computer program product comprising a computer program, characterized in that When the above computer program is executed by a processor, A lyrics progression control method according to any one of claims 6 to 10 is implemented.
Citation Information
Patent Citations
Automatic musical accompaniment apparatus
JP2005266080A
Device and program for synthesizing singing
JP2014010190A
Singing voice output control device
JP2017003625A