Electronic musical instrument, electronic musical instrument control method, and program
By configuring pitch designation units and performance style output units in electronic musical instruments and synthesizing appropriate musical sound data using trained acoustic models, the problem of insufficient expressiveness caused by real-time performance speed changes is solved, and effective musical performance at different performance speeds is achieved.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- CASIO COMPUTER CO LTD
- Filing Date
- 2021-08-13
- Publication Date
- 2026-08-04
AI Technical Summary
Existing electronic musical instruments cannot accurately predict the appropriate sound waveform to match the changes in playing speed between notes during real-time performance. This results in insufficient expressiveness when playing slowly or a slow rise in sound waveform when playing fast, making it difficult to achieve effective musical expression.
It is equipped with a pitch specification unit, a performance style output unit, and a sound generation model unit. By inputting trained acoustic model parameters, it synthesizes and outputs musical sound data corresponding to the pitch data and performance style data during performance, thereby achieving adaptation to changes in real-time performance speed.
It can accurately infer and output appropriate sound waveforms that match the playing speed of notes that change in real time, improving the expressiveness and clarity of electronic instruments at different playing speeds.
Smart Images

Figure CN116057624B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to electronic musical instruments, electronic musical instrument control methods and programs, which are used to output speech sounds by driving a trained acoustic model in response to operations on operating elements such as a keyboard. Background Technology
[0002] In electronic musical instruments, to complement the expressiveness of vocal voice and live instruments, which are weaknesses of the pulse code modulation (PCM) method in related technologies, a technique for training an acoustic model based on actual performance operations has been designed and put into practical use. In this model, the human vocalization mechanism and the sound generation mechanism of the instrument are modeled by performing digital signal processing through machine learning based on singing and playing operations, and by driving the trained acoustic model to infer and output the sound waveform data of the singing voice or musical sound (e.g., Patent Document 1).
[0003] Citation List
[0004] Patent documents
[0005] Patent Document 1: Japanese Patent No. 6610714 Summary of the Invention
[0006] Technical issues
[0007] For example, when machine learning is used to generate vocal or musical waveforms, the generated waveforms typically change depending on the rhythm, phrasing, and performance style. For instance, the duration of consonant sounds in the human voice, the duration of wind instrument sounds, and the duration of noise components at the beginning of string playing on bowed instruments are longer in slow, low-paced performances with few notes, resulting in a highly expressive and lively sound, while they are shorter in fast-paced performances with many notes, resulting in an articulated sound.
[0008] However, when users perform live on keyboards or similar devices, the acoustic model cannot convey the changing tempos between notes in response to changes in the score division for each note or differences in the musical phrases played in the sound source device. This prevents the acoustic model from inferring the appropriate sound waveform corresponding to the changes in tempo between notes. As a result, for example, slow playing lacks expressiveness, or conversely, fast playing produces a slow rise in the sound waveform, making it difficult to play.
[0009] Therefore, the object of the present invention is to enable the deduction of an appropriate sound waveform that matches the change in playing speed between notes that change in real time.
[0010] Solution to the problem
[0011] An example electronic musical instrument includes: a pitch designation unit configured to output performance pitch data specified during performance; a performance style output unit configured to output performance style data indicating the performance style during performance; and a sound generation model unit configured to synthesize and output musical sound data corresponding to the performance pitch data and the performance style data during performance based on acoustic model parameters inferred by inputting the performance pitch data and the performance style data into a trained acoustic model.
[0012] Another example of this aspect is an electronic musical instrument comprising: a lyrics output unit configured to output performance-time lyrics data indicating the lyrics during performance; a pitch designation unit configured to output performance-time pitch data tuned to the output of the lyrics during performance; a performance style output unit configured to output performance-time performance style data indicating the performance style during performance; and a vocal model unit configured to synthesize and output performance-time singing voice data corresponding to the performance-time lyrics data, the performance-time pitch data, and the performance-time performance style data based on acoustic model parameters inferred by inputting the performance-time lyrics data, the performance-time pitch data, and the performance-time performance style data into a trained acoustic model.
[0013] Beneficial effects of the present invention
[0014] According to the present invention, it is possible to deduce an appropriate speech waveform that matches the change in playing speed between notes that change in real time. Attached Figure Description
[0015] Figure 1 An example of the appearance of an embodiment of an electronic keyboard musical instrument is shown.
[0016] Figure 2 This is a block diagram illustrating an example of the hardware configuration of a control system for an electronic keyboard musical instrument.
[0017] Figure 3 This is a block diagram showing an example configuration of the speech training unit and the speech synthesis unit.
[0018] Figure 4A This is an explanatory diagram showing an example of musical notation division that forms the basis of singing style.
[0019] Figure 4B This is an explanatory diagram showing an example of musical notation division that forms the basis of singing style.
[0020] Figure 5A This demonstrates how differences in the rhythm of the performance cause changes in the waveform of the singing voice.
[0021] Figure 5B This demonstrates how differences in the rhythm of the performance cause changes in the waveform of the singing voice.
[0022] Figure 6 This is a block diagram showing an example configuration of the lyrics output unit, pitch specification unit, and performance style output unit.
[0023] Figure 7 An example of data configuration for this embodiment is shown.
[0024] Figure 8 This is a main flowchart illustrating an example of the control processing of an electronic musical instrument in this embodiment.
[0025] Figure 9A This is a flowchart showing a detailed example of the initialization process.
[0026] Figure 9B This is a flowchart illustrating a detailed example of rhythm change processing.
[0027] Figure 9C This is a flowchart showing a detailed example of how song processing begins.
[0028] Figure 10 This is a flowchart illustrating a detailed example of switch processing.
[0029] Figure 11 This is a flowchart illustrating a detailed example of keyboard processing.
[0030] Figure 12 This is a flowchart illustrating a detailed example of automatic performance interruption handling.
[0031] Figure 13 This is a flowchart illustrating a detailed example of song playback processing. Specific Implementation
[0032] In the following, embodiments of the present invention will be described in detail with reference to the accompanying drawings.
[0033] Figure 1An example of the appearance of an embodiment of the electronic keyboard musical instrument 100 is shown. The electronic keyboard musical instrument 100 includes a keyboard 101 consisting of a plurality of keys as operating elements; a first switch panel 102 configured to indicate various settings such as volume setting, song playback tempo setting (described later), performance tempo mode setting (described later), performance tempo adjustment setting (described later), song playback start (described later), and accompaniment playback start (described later); a second switch panel 103 configured to select a song or accompaniment and timbre; a liquid crystal display (LCD) 104 configured to display sheet music and lyrics (described later) during song playback; and information related to the various settings. Furthermore, although not specifically shown, the electronic keyboard musical instrument 100 includes speakers configured to emit musical sounds generated by performance and provided on a rear portion, side portion, rear surface portion, etc.
[0034] Figure 2 It shows Figure 1 This is an example of the hardware configuration of an embodiment of the control system 200 for the electronic keyboard musical instrument 100 shown. Figure 2 In the control system 200, there are CPU (Central Processing Unit) 201, ROM (Read-Only Memory) 202, RAM (Random Access Memory) 203, sound source LSI (Large-Scale Integration) 204, speech synthesizer LSI 205, and... Figure 1 The keyboard 101, the first switch panel 102, and the second switch panel 103 shown are connected to a key scanner 206 and a key scanner 206. Figure 1 The LCD controller 208 connected to the shown LCD 104 and the network interface 219 configured to send MIDI data to and receive MIDI data from an external network are respectively connected to the system bus 209. Furthermore, a timer 210 for controlling automatic playing sequences is connected to the CPU 201. Additionally, music sound data 218 and singing sound data 217 output from the sound source LSI 204 and speech synthesis LSI 205, respectively, are converted into analog music sound output signals and analog singing sound output signals by D / A converters 211 and 212. The analog music sound output signals and analog singing sound output signals are mixed in mixer 213, and the mixed signal is amplified in amplifier 214 before being output from a speaker or output terminal (not specifically shown).
[0035] CPU 201 is configured to execute a control program loaded from ROM 202 into RAM 203 while using RAM 203 as working memory. Figure 1 The control operation of the electronic keyboard instrument 100 is shown. In addition, the ROM 202 (non-temporary recording medium) is configured to store music segment data, including lyrics data and accompaniment data, in addition to the control program and various types of fixed data.
[0036] The timer 210 used in this embodiment is implemented on the CPU 201 and is configured, for example, to count the automatic playing in the electronic keyboard instrument 100.
[0037] The sound source LSI 204 is configured to read music sound waveform data from, for example, a waveform ROM (not specifically shown), and output it as music sound data 218 to the D / A converter 211 in response to sound generation control data 216 from the CPU 201. The sound source LSI 204 is capable of producing 256-voice polyphony.
[0038] When given a speech synthesis LSI 205, the speech synthesis LSI synthesizes singing voice data 217 corresponding to the performance singing voice data 215 from the CPU 201, the text data of the lyrics (performance lyrics data), the data specifying each pitch corresponding to each lyric (performance pitch data), and the data about how to sing (performance style data), and outputs the singing voice data to the D / A converter 212.
[0039] Key scanner 206 is configured to scan periodically. Figure 1 The key press / release state on the keyboard 101 shown, as well as the switch operation state of the first switch panel 102 and the second switch panel 103, are transmitted to the CPU 201 via an interrupt to transmit the state change.
[0040] LCD controller 208 is an IC (integrated circuit) configured to control the display state of LCD 104.
[0041] Figure 3 This is a block diagram illustrating an example configuration of the speech synthesis unit and the speech training unit in this embodiment. Here, the speech synthesis unit 302 is built into the electronic keyboard instrument 100, serving as a component of the speech synthesis unit. Figure 2 This is a function performed by the LSI 205 speech synthesis system.
[0042] The speech synthesis unit 302 processes the automatic playback of lyrics (hereinafter referred to as "song playback") Figure 1 The key presses on the keyboard 101, via Figure 2The key scanner 206 in the speech synthesis unit 302 takes into account the performance-time singing voice data 215, which includes lyrics, pitch, and information about how to sing, as instructed by the CPU 201, and synthesizes and outputs singing voice data 217, which will be described later. At this time, the processor of the speech synthesis unit 302 performs voice processing, inputting the performance-time singing voice data 215, which includes lyrics information generated by the CPU 201 in response to the operation of any one of the multiple keys (operating elements) on the keyboard 101, pitch information associated with any key, and information about how to sing, into the performance-time singing voice analysis unit 307, inputting the performance-time language feature sequence 316 output from the performance-time singing voice analysis unit into the trained acoustic model stored in the acoustic model unit 306, and outputting singing voice data 217, which infers the singer's singing voice based on the spectrum information 318 and sound source information 319 output as a result by the acoustic model unit 306.
[0043] For example, such as Figure 3 As shown, the speech training unit 301 can be implemented by... Figure 1 The electronic keyboard instrument 100 in the instrument is a separate function that exists on an external server computer 300. Alternatively, although not in... Figure 3 The text shows that if Figure 2 If the speech synthesis LSI 205 has idle processing capabilities, then the speech training unit 301 can also be built into the electronic keyboard instrument 100 as a function performed by the speech synthesis LSI 205.
[0044] Figure 2 The speech training unit 301 and speech synthesis unit 302 shown are implemented, for example, based on the "statistical parametric speech synthesis based on deep learning" technique described in Non-Patent Document 1 cited below.
[0045] (Non-patent literature 1)
[0046] Kei Hashimoto and Shinji Takaki, “Statistical parametric speech synthesis based on deep learning,” Journal of the Acoustical Society of Japan, Vol. 73, No. 1 (2017), pp. 55-62.
[0047] Figure 2The speech training unit 301 is composed of Figure 3 The functions performed by the external server computer 300 shown include training a singing voice analysis unit 303, training an acoustic feature extraction unit 304, and a model training unit 305.
[0048] The voice training unit 301 uses, for example, voice recordings of a singer performing multiple songs in an appropriate genre as training singing voice data 312. Furthermore, it prepares text data of the lyrics for each song (training lyrics data), data specifying each pitch corresponding to each lyric (training pitch data), and data indicating the singing style of the training singing voice data 312 (training performance style data) as training singing voice data 311. As training performance style data, the training pitch data is measured sequentially at sequentially specified time intervals, and each data point indicating the sequentially measured time intervals is specified.
[0049] Training singing voice data 311, including training lyrics data, training pitch data, and training performance style data, is input into training singing voice analysis unit 303. Training singing voice analysis unit 303 analyzes the input data. Therefore, training singing voice analysis unit 303 estimates and outputs training language feature sequence 313, which is a discrete digital sequence representing the phonemes, pitches, and singing styles corresponding to training singing voice data 311.
[0050] In response to the input of training singing voice data 311, the training acoustic feature extraction unit 304 receives and analyzes training singing voice data 312 recorded via a microphone or the like when a specific singer sings lyrics corresponding to the training singing voice data 311. Therefore, the training acoustic feature extraction unit 304 extracts a training acoustic feature sequence 314 representing the features of the speech sounds corresponding to the training singing voice data 312 and outputs it as teacher data.
[0051] The training language feature sequence 313 is represented by the following symbols.
[0052] [Expression 1]
[0053] l
[0054] The acoustic model is represented by the following symbols.
[0055] [Expression 2]
[0056] λ
[0057] The training acoustic feature sequence 314 is represented by the following symbols.
[0058] [Expression 3]
[0059] O
[0060] The probability of generating training acoustic feature sequence 314 is represented by the following symbol.
[0061] [Expression 4]
[0062] P(o|l,λ)
[0063] The acoustic model that maximizes the probability of generating the training acoustic feature sequence 314 is represented by the following notation.
[0064] [Expression 5]
[0065]
[0066] The model training unit 305 estimates the acoustic model by performing machine learning from the training language feature sequence 314 and the acoustic model according to the following equation (1), which maximizes the probability of generating the training acoustic feature sequence 314. That is, the relationship between the language feature sequence as text and the acoustic feature sequence as speech sound is represented by a statistical model called the acoustic model.
[0067] [Expression 6]
[0068]
[0069] Here, the following symbols indicate the calculation of the value of the independent variable below the symbol, which gives the maximum value of the function to the right of the symbol.
[0070] [Expression 7]
[0071] arg max
[0072] The model training unit 305 outputs training result data 315, which represents the acoustic model calculated as a result of machine learning through the calculation shown in equation (1). The calculated acoustic model is represented by the following symbols.
[0073] [Expression 8]
[0074]
[0075] like Figure 3 As shown, for example, in Figure 1 When the electronic keyboard instrument 100 is manufactured, the training result data 315 can be stored in the electronic keyboard instrument 100. Figure 2 The control system ROM 202 shown can be accessed from the electronic keyboard instrument 100 when it is powered on. Figure 2 The ROM 202 in the speech synthesis LSI 205 is loaded into the acoustic model unit 306, which will be described later. Alternatively, for example, as Figure 3As shown, the training result data 315 can also be downloaded via network interface 219 from a network such as the Internet and a network such as a USB (Universal Serial Bus) cable (not specifically shown) to the acoustic model unit 306 in the speech synthesis LSI 205 (described later) through network interface 219, via user operation on the second switch panel 103 of the electronic keyboard instrument 100. Alternatively, in addition to the speech synthesis LSI 205, the trained acoustic model can also be implemented in hardware using an FPGA (Field-Programmable Gate Array) or similar device, and then used as an acoustic model unit.
[0076] The speech synthesis unit 302, which performs the functions to be executed by the speech synthesis LSI 205, includes a singing voice analysis unit 307, an acoustic model unit 306, and a vocalization model unit 308. The speech synthesis unit 302 performs statistical speech synthesis processing, predicts using a statistical model called the acoustic model set in the acoustic model unit 306, and sequentially synthesizes and outputs singing voice data 217 corresponding to the singing voice data 215 that are sequentially input during performance.
[0077] As a result of the tuning between the user's performance and the automatic performance, the singing voice data 215 during performance is input to the singing voice analysis unit 307 during performance. This singing voice data 215 includes data from... Figure 2 The CPU 201 specifies information regarding performance-time lyrics data (phonemes corresponding to the lyrics text), performance-time pitch data, and performance-time style data (data about how to sing), and the performance-time singing voice analysis unit 307 analyzes the input data. Therefore, the performance-time singing voice analysis unit 307 analyzes and outputs a performance-time language feature sequence 316 representing the phonemes, parts of speech, words, pitch, and singing style corresponding to the performance-time singing voice data 215.
[0078] In response to the input of the performance-time speech feature sequence 316, the acoustic model unit 306 estimates and outputs the performance-time acoustic feature sequence 317, which consists of acoustic model parameters corresponding to the input performance-time speech feature sequence. The performance-time speech feature sequence 316 input from the performance-time singing voice analysis unit 307 is represented by the following symbols.
[0079] [Expression 9]
[0080] l
[0081] The acoustic model set as training result data 315 by machine learning in model training unit 305 is represented by the following symbols.
[0082] [Expression 10]
[0083]
[0084] The acoustic feature sequence 317 during performance is represented by the following symbols.
[0085] [Expression 11]
[0086] o
[0087] The probability of generating the acoustic feature sequence 317 during performance is represented by the following symbols.
[0088] [Expression 12]
[0089]
[0090] The estimated value of the acoustic feature sequence 317 during performance is the acoustic model parameter that maximizes the probability of generating the acoustic feature sequence 317 during performance, and is represented by the following symbols.
[0091] [Expression 13]
[0092]
[0093] The acoustic model unit 306 estimates the estimated value of the acoustic feature sequence 317 during performance based on the performance language feature sequence 316 input from the performance singing voice analysis unit 307 and the acoustic model set as training result data 315 in the model training unit 305 through machine learning, according to the following equation (2). The estimated value of the acoustic feature sequence 317 during performance is the acoustic model parameter that maximizes the probability that the acoustic feature sequence 317 during performance will be generated.
[0094] [Expression 14]
[0095]
[0096] In response to the input of the acoustic feature sequence 317, the vocal model unit 308 synthesizes and outputs singing voice data 217 corresponding to the singing voice data 215 specified from the CPU 201 during performance. This singing voice data 217 is transmitted via mixer 213 and amplifier 214 from... Figure 2 The output of the D / A converter 212 in the middle is emitted by a speaker that is never specifically shown.
[0097] The acoustic features represented by the training acoustic feature sequence 314 or the performance acoustic feature sequence 317 include spectral information modeling the human vocal tract and sound source information modeling the human vocal cords. As spectral information (parameters), for example, Mel cepstrum, line spectrum pairs (LSPs), etc., can be used. As sound source information, power values and the fundamental frequency (F0) indicating the pitch frequency of human speech can be used. The vocal model unit 308 includes a sound source generation unit 309 and a synthesis filter unit 310. The sound source generation unit 309 is a unit that models the human vocal cords and, in response to the sequence of sound source information 319 sequentially input from the acoustic model unit 306, generates sound source signal data consisting of pulse sequence data (in the case of voiced phonemes) that periodically repeats at the fundamental frequency (F0) and the power values contained in the sound source information 319, such as white noise data (in the case of unvoiced phonemes) with power values contained in the sound source information 319, or a mixture thereof. The synthesis filter unit 310 is a unit that models the human vocal tract and forms a digital filter that models the vocal tract based on the sequence of spectral information 318 sequentially input from the acoustic model unit 306. It generates and outputs singing sound data 321, which is digital signal data, by using sound source data input from the sound source generation unit 309 as excitation source signal data.
[0098] The sampling frequency of the training vocal data 312 and vocal data 217 is, for example, 16 kHz. For example, when using Mel-Cepstral parameters obtained through Mel-Cepstral Analysis for the spectral parameters included in the training acoustic feature sequence 314 and the performance acoustic feature sequence 317, the frame update period is, for example, 6 msec. Furthermore, during Mel-Cepstral Analysis, the analysis window length is 25 msec, the window function is a Blackman window function, and the analysis order is 24.
[0099] As by Figure 3 The specific processing of statistical speech synthesis performed by the speech training unit 301 and speech synthesis unit 302 in the acoustic model unit 306, which represents the acoustic model based on the training result data 315 set in the acoustic model unit 306, can employ a method using a Hidden Markov Model (HMM) or a method using a Deep Neural Network (DNN). Since specific embodiments of this are disclosed in the aforementioned Patent Document 1, detailed descriptions are omitted in this application.
[0100] By Figure 3The statistical speech synthesis processing performed by the speech training unit 301 and speech synthesis unit 302 shown realizes the electronic keyboard instrument 100, which outputs singing voice data 217 of a specific performer singing well by allowing singing voice data 215 during performance to be sequentially input into the acoustic model unit 306. The singing voice data 215 during performance includes the lyrics and pitch of the song played by the user pressing the key. The acoustic model unit 306 is equipped with an acoustic model that has been trained on the singing voice of a specific singer.
[0101] Here, it is normal for there to be differences in the singing style between the fast and slow sections of the melody. Figure 4A and Figure 4B This is an explanatory diagram showing an example of musical score division that forms the basis of singing methods. Figure 4A An example of the musical score for the lyrics melody of a fast passage is shown, while Figure 4B Examples of musical scores for slow-tempo lyrics are shown. In these examples, the pitch-changing patterns are similar. However, Figure 4A The musical notation for a sequence of sixteenth notes (notes that are 1 / 4 the length of a quarter note) is shown, while Figure 4B The musical notation for a sequence of quarter notes is shown. Therefore, regarding the tempo of changing pitch, Figure 4A The speed at which musical scores are divided is Figure 4B The tempo of a musical score is four times that of a given piece. In fast passages, the consonant parts of the vocal register cannot be sung (played) well unless shortened. Conversely, in fast passages, when the consonant parts of the vocal register are lengthened, a highly expressive singing (playing) can be achieved. As mentioned above, even when the pitch change pattern is the same, the difference in the length of each note (quarter note, eighth note, sixteenth note, etc.) in a singing melody will lead to a difference in singing (playing) tempo. However, needless to say, even when singing (playing) the exact same score, the playing tempo will differ when the rhythm changes. In the following description, the time interval between notes (sound generation speed) caused by the above two factors will be described as "playing rhythm" to distinguish it from the rhythm of a regular song.
[0102] Figure 5A and Figure 5B It is shown as follows Figure 4A and Figure 4B The diagram shows the changes in the waveform of the singing voice caused by differences in performance rhythm. Example: Figure 5A and Figure 5BThe image shows an example waveform of a sung sound when the / ga / phonological sound is produced. The / ga / phonological sound is a combination of the consonant / g / and the vowel / a / . In many cases, the duration (time length) of the consonant portion is typically from tens of milliseconds to approximately 200 milliseconds. Here, Figure 5A An example of the vocal waveform when singing a fast passage is shown. Figure 5B An example of the vocal waveform when singing a slow passage is shown. Figure 5A and Figure 5B The difference between the waveforms lies in the length of the consonant / g / . It can be seen that when singing in a fast passage, such as... Figure 5A As shown, the duration of consonant sounds is shorter; conversely, when singing slow passages, such as... Figure 5B As shown, the consonant part is pronounced for a longer duration. When singing in fast passages, priority is given to the initial speed of the vocalization, and it is not necessary to pronounce the consonants clearly. However, when singing in slow passages, the consonants are often pronounced longer and more clearly, which increases the clarity of the words.
[0103] In order to reflect the differences in performance rhythm as described above in the changes in singing voice data, in the process of... Figure 3 In the statistical speech synthesis processing performed by the speech training unit 301 and speech synthesis unit 302 shown, the training singing voice data 311 input to the speech training unit 301 is supplemented with training lyrics data indicating lyrics, training pitch data indicating pitch, and training performance style data indicating singing style. Information about performance rhythm is also included in the training performance style data. The training singing voice analysis unit 303 in the speech training unit 301 analyzes the training singing voice data 311 to generate a training language feature sequence 313. The model training unit 305 in the speech training unit 301 performs machine learning using the training language feature sequence 313. As a result, the model training unit 305 can output a trained acoustic model including information about performance rhythm as training result data 315, and store it in the acoustic model unit 306 in the speech synthesis unit 302 of the speech synthesis LSI 205. As training performance style data, training pitch data is measured sequentially at sequentially specified time intervals, and each performance rhythm data indicating the sequentially measured time intervals is specified. In this way, the model training unit 305 of this embodiment can perform training that can derive a trained acoustic model, incorporating differences in performance rhythm due to singing style.
[0104] On the other hand, in the speech synthesis unit 302, which includes an acoustic model unit 306 with a trained acoustic model configured as described above, performance style data indicating the singing style is added to performance lyrics data indicating the lyrics, and performance pitch data indicating the pitch is added to performance singing voice data 215. Information about the performance rhythm can also be included in the performance style data. The performance singing voice analysis unit 307 in the speech synthesis unit 302 analyzes the performance singing voice data 215 to generate a performance language feature sequence 316. Then, the acoustic model unit 306 in the speech synthesis unit 302 outputs corresponding spectral information 318 and sound source information 319 by inputting the performance language feature sequence 316 into the trained acoustic model, and provides the spectral information and sound source information to the synthesis filter unit 310 and sound source generation unit 309 in the vocal model unit 308, respectively. As a result, the vocal model unit 308 can output singing voice data 217, wherein the differences in performance rhythm caused by the singing style result in... Figure 5A and Figure 5B The changes in the length of consonants, etc., are already reflected. In other words, appropriate singing voice data 217 can be deduced to match the changes in playing speed between notes that change in real time.
[0105] Figure 6 This is a block diagram showing an example configuration of the lyrics output unit, pitch specification unit, and performance style output unit, implemented as follows: Figures 8 to 11 The flowchart shown is composed of Figure 2 The control processing performed by the CPU 201 shown above to generate the singing sound data 215 during the performance described above (described later) is the process.
[0106] Lyrics output unit 601 outputs lyrics data 609 for each performance moment, indicating the lyrics during performance, and includes it in the output to... Figure 2 The speech synthesis LSI 205 in the speech synthesis LSI 205 contains singing voice data 215 for each performance. Specifically, the lyrics output unit 601 sequentially reads each timing data 605 in the music segment data 604 for song playback that is pre-loaded from ROM 202 into RAM 203 by CPU 201, and reads each lyric data (lyric text) 608 in each event data 606 stored in pairs as music segment data 604 according to the timing sequence indicated by each timing data 605, and sets them as performance lyrics data 609 respectively.
[0107] The pitch designation unit 602 outputs pitch data 610 for each pitch tuned to the output of each lyric during performance, and includes it in the output to... Figure 2The speech synthesis LSI 205 in the speech synthesis LSI 205 performs each singing voice data 215. Specifically, the pitch specification unit 602 sequentially reads each timing data 605 in the music segment data 604 loaded into RAM 203 for song playback, and when it is pressed by the user... Figure 1 When any key on the keyboard 101 is pressed, the pitch information associated with that key is input via the key scanner 206 at the timing indicated by each timing data 605, and the pitch information is set to the performance pitch data 610. Furthermore, if the user does not press any key at the timing indicated by each timing data 605... Figure 1 When any key is pressed on the keyboard 101, the pitch designation unit 602 sets the pitch data 607 of the event data 606 of the music segment data 604, which is stored in pairs with the timing data 605, as the performance pitch data 610.
[0108] The performance style output unit 603 outputs performance style data 611, which is the singing style of the performance during the performance, and includes it in the output to... Figure 2 The singing voice data 215 for each performance of the LSI 205 speech synthesis system.
[0109] Specifically, when the user is Figure 1 When the performance rhythm mode is set to free mode on the first switch panel 102, as will be described below, the performance style output unit 603 sequentially measures the time interval of the pitch specified by the user's key presses during performance, and sets each performance rhythm data indicating the sequentially measured time interval as performance style data 611 for each performance.
[0110] On the other hand, when the user is not in Figure 1 When the performance rhythm mode is set to free mode on the first switch panel 102, as will be described below, the performance style output unit 603 sets each performance rhythm data to performance style data 611 for each performance, each performance rhythm data corresponding to each time interval indicated by each timing data 605 read sequentially from the music segment data 604 loaded into RAM 203 for song playback.
[0111] Additionally, when users are Figure 1 When the performance rhythm mode is set to the performance rhythm adjustment mode on the first switch panel 102 for intentional change of the performance rhythm mode, as will be described below, the performance style output unit 603 intentionally changes the value of each performance rhythm data obtained in sequence as described, based on the value of the performance rhythm adjustment setting, and sets each changed performance rhythm data as the performance style data 611 during performance.
[0112] In this way, by Figure 1 The CPU 201 executes the lyrics output unit 601, pitch specification unit 602, and performance style output unit 603. Each function generates performance-time singing sound data 215 at a timed point when a key press occurs during song playback, based on a user's key press or a key press event. This data includes performance-time lyrics data 609, performance-time pitch data 610, and performance-time performance style data 611. This data can be sent to a system with... Figure 2 or Figure 3 The speech synthesis unit 302 in the LSI 205 configured in the speech synthesis system.
[0113] The following will describe the usage in detail. Figures 3 to 6 The statistical speech synthesis processing described in [the text] Figure 1 and Figure 2 Operation of an embodiment of the electronic keyboard instrument 100. Figure 7 This is shown in this embodiment from Figure 2 The diagram illustrates a detailed data configuration example of music segment data loaded from ROM 202 into RAM 203. This data configuration example conforms to the standard MIDI file format, one of the MIDI (Musical Instrument Digital Interface) file formats. The music segment data is configured using data blocks called chunks. Specifically, the music segment data is configured as follows: a head chunk at the beginning of the file, a first track chunk following the head chunk and storing the lyrics data, and a second track chunk storing the performance data for the accompaniment.
[0114] The header chunk consists of four values: ChunkID, ChunkSize, Format Type, NumberOfTrack, and TimeDivision. ChunkID is a 4-byte ASCII code "4D 54 68 64" (hexadecimal) corresponding to the four-width characters "MThd", indicating that this chunk is a header chunk. ChunkSize is 4 bytes of data, representing the data length of the FormatType, NumberOfTrack, and TimeDivision portions of the header chunk, excluding ChunkID and ChunkSize. The data length is fixed at six bytes "00 00 00 06" (hexadecimal). In this embodiment, FormatType is 2 bytes "00 01" (hexadecimal), meaning the format type is format 1, which uses multiple tracks. In this embodiment, NumberOfTrack is 2 bytes "00 02" (hexadecimal), indicating the use of two tracks corresponding to the lyrics and accompaniment sections. TimeDivision is data indicating the time base value, which indicates the resolution of each quarter note. In this embodiment, it is a 2-byte data "01E0" in decimal 480 (the number is in hexadecimal).
[0115] The first track block indicates the lyrics section, which corresponds to Figure 6 The music segment data 604 is configured with the following items: ChunkID, ChunkSize, and the corresponding... Figure 6 The timing data 605's DeltaTime_1[i] and the corresponding Figure 6 The performance data pair consists of Event_1[i] of event data 606 in the middle (0≤i≤L-1). In addition, the second track block corresponds to the accompaniment part, which is configured with the following items: ChunkID, ChunkSize, and performance data pair consisting of DeltaTime_2[i] as the timing data of the accompaniment part and Event_2[j] as the event data of the accompaniment part (0≤j≤M-1).
[0116] Each ChunkID in the first and second track blocks is a 4-byte ASCII code "4D 54 72 6B" (in hexadecimal) corresponding to the four half-width characters "MTrk", indicating that the block is a track block. Each ChunkSize in the first and second track blocks is 4 bytes of data indicating the data length of each track block, excluding ChunkID and ChunkSize.
[0117] DeltaTime_1[i], i.e. Figure 6The timing data 605 is a variable-length data of 1 to 4 bytes indicating the waiting time (relative time) from the execution time of Event_1[i-1]. Event_1[i-1] is the data immediately preceding it. Figure 6 Event data 605. Similarly, DeltaTime_2[i], which is the timing data for the accompaniment section, is a variable-length data of 1 to 4 bytes indicating the waiting time (relative time) starting from the execution time of Event_2[i-1], which is the event data of the accompaniment section immediately preceding it.
[0118] Event_1[i], i.e. Figure 6 Event data 606 in the first track block / lyrics section of this embodiment is a meta-event with two pieces of information, namely the spoken text of the lyrics and the pitch. Event_2[i], which is the event data of the accompaniment section, is a MIDI event that specifies the note-on or note-off of the accompaniment sound in the second track block / accompaniment section, or a meta-event that specifies the rhythm of the accompaniment sound.
[0119] In each performance data pair DeltaTime_1[i] and Event_1[i] of the first track block / lyrics section, Event_1[i], which is event data 606, is executed after waiting for DeltaTime_1[i] from the execution time of Event_1[i-1]. DeltaTime_1[i] is timing data 605, and Event_1[i-1] is the event data 606 immediately preceding it. This allows the song playback to proceed. On the other hand, in each performance data pair DeltaTime_2[i] and Event_2[i] of the second track block / accompaniment section, Event_2[i], which is event data, is executed after waiting for DeltaTime_2[i] from the execution time of Event_2[i-1]. DeltaTime_2[i] is timing data, and Event_2[i-1] is the event data immediately preceding it. This allows the automatic accompaniment to progress.
[0120] Figure 8 This is a main flowchart illustrating an example of the control processing of an electronic musical instrument according to this embodiment. For this control processing, for example... Figure 2 The CPU 201 executes the control processing program loaded from ROM 202 into RAM 203.
[0121] After the initialization process (step S801) is executed first, the CPU 201 repeats a series of processes from step S802 to step S808.
[0122] In this repetitive process, CPU 201 first executes the switching process (step S802). Here, CPU 201 is based on the data from... Figure 2 The key scanner 206 in the middle is interrupted to perform with Figure 1 The corresponding processing for switch operations on the first switch panel 102 or the second switch panel 103 will be referred to later. Figure 10 The flowchart in the document describes the switch processing in detail.
[0123] Next, CPU 201 is based on... Figure 2 The key scanner 206 interrupts to perform the determination of whether an operation has been performed. Figure 1 The keyboard processing of any key on keyboard 101 continues accordingly (step S803). During keyboard processing, in response to a user operation of pressing or releasing any key, CPU 201 outputs an indication. Figure 2 The sound source LSI 204 in the middle starts or stops generating sound, which is music sound control data 216. Additionally, in keyboard processing, the CPU 201 performs calculations to treat the time interval from the immediately preceding key press to the current key press as performance rhythm data. (See later...) Figure 11 The flowchart in the document describes keyboard handling in detail.
[0124] Next, CPU 201 will process... Figure 1 The data displayed on LCD 104, and via Figure 2 The LCD controller 208 performs display processing (step S804) to display data on the LCD 104. Examples of data to be displayed on the LCD 104 include lyrics corresponding to the singing voice data 217 being played, musical scores corresponding to the melody and accompaniment of the lyrics, and information related to various settings.
[0125] Next, CPU 201 executes song playback processing (step S805). In song playback processing, CPU 201 generates and sends performance-time singing voice data 215 to speech synthesis LSI 205. This performance-time singing voice data 215 includes lyrics, pitch, and rhythm for manipulating speech synthesis LSI 205 based on song playback. (See later...) Figure 13 The flowchart in the document describes the song playback process in detail.
[0126] Subsequently, CPU 201 performs sound source processing (step S806). In sound source processing, CPU 201 performs control processing, such as processing for controlling the envelope of the musical sound generated in sound source LSI 204.
[0127] Subsequently, the CPU 201 performs speech synthesis processing (step S807). In the speech synthesis processing, the CPU 201 controls the execution of speech synthesis by the speech synthesis LSI 205.
[0128] Finally, CPU 201 determines whether the user has pressed the power off switch (not specifically shown) to turn off the power (step S808). If the determination in step S808 is "No", CPU 201 returns to the processing in step S802. If the determination in step S808 is "Yes", CPU 201 ends. Figure 8 The control process is shown in the flowchart, and the power to the electronic keyboard instrument 100 is turned off.
[0129] Figure 9A , Figure 9B and Figure 9C Each is shown in Figure 8 During the switching process of step S802 Figure 8 The initialization process in step S801, Figure 10 The rhythm change processing in step S1002 and similarly Figure 10 A detailed flowchart of a step S1006, which involves starting song processing, is provided below.
[0130] First of all, Figure 9A It shows Figure 8 A detailed example of the initialization process in step S801 is that the CPU 201 executes the TickTime initialization process. In this embodiment, the lyrics and automatic accompaniment are performed in a time unit called TickTime. Figure 7 The time base value of the TimeDivision value specified in the header block of the music segment data indicates the resolution per quarter note. For example, if this value is 480, then the length of each quarter note is 480 TickTimes. The DeltaTime_1[i] and DeltaTime_2[i] values indicate... Figure 7 The waiting time in the track blocks of the music segment data is also calculated in units of TickTime. Here, the actual number of seconds corresponding to 1 TickTime varies depending on the tempo specified for the music segment data. Using the tempo value as Tempo (beats per minute) and the time base value as TimeDivision, the number of seconds per unit of TickTime is calculated using the following equation (3).
[0131] [Expression 15]
[0132] TickTime[sec]=60 / Tempo / TimeDivision (3)
[0133] Therefore, in Figure 9A In the initialization process shown in the flowchart, CPU 201 first calculates TickTime (sec) through arithmetic operations corresponding to equation (10) (step S901). Note that it is assumed that the specified value of tempo, such as 60 (beats per second), is stored in the initial state. Figure 2 In ROM 202. Alternatively, the rhythm value at the end of the previous processing can be stored in non-volatile memory.
[0134] Next, CPU 201 uses the TickTime (sec) calculated in step S901 to... Figure 2 Timer 210 in the CPU is set to trigger a timer interrupt (step S902). As a result, whenever TickTime (sec) has elapsed, timer 210 sends an interrupt (hereinafter referred to as "automatic performance interrupt") to CPU 201 for song playback and automatic accompaniment. Therefore, the CPU 201 performs automatic performance interrupt processing based on the automatic performance interrupt. Figure 12 In the (described later) section, each TickTime executes control processing for song playback and automatic accompaniment.
[0135] Subsequently, CPU 201 performs additional initialization processes, such as those for initialization. Figure 2 Processing of RAM 203 in the CPU (step S903). After this, CPU 201 terminates. Figure 8 The initialization process in step S801, such as... Figure 9A The flowchart shown.
[0136] A description will follow later. Figure 9B and Figure 9C The flowchart in the document. Figure 10 It is shown Figure 8 A flowchart illustrating a detailed example of the switching process in step S802.
[0137] The CPU 201 first determines whether the rhythm of the lyrics and automatic playback has been changed via the rhythm change switch on the first switch panel 102 (step S1001). When it is determined to be "yes", the CPU 201 executes the rhythm change processing (step S1002). See below for further details. Figure 9B The process is described in detail. When the determination in step S1001 is "No", CPU 201 skips the process in step S1002.
[0138] Next, CPU 201 determines whether it has been used. Figure 1The second switch panel 103 selects any song (step S1003). When "yes" is determined, the CPU 201 executes song loading processing (step S1004). This processing loads the song with... Figure 7 Music segment data in the data structure described in [the document] is loaded from ROM 202. Figure 2 The processing is performed in RAM 203. It's important to note that song loading processing can be performed before the performance begins, rather than during the performance itself. Figure 7 Subsequent data access to the first or second track block in the data structure shown is performed on the music segment data loaded into RAM 203. When the determination in step S1003 is "No", CPU 201 skips the processing in step S1004.
[0139] Subsequently, CPU 201 determines whether... Figure 1 The song start switch has been activated on the first switch panel 102 (step S1005). When the setting is confirmed as "yes", the CPU 201 executes the song start process (step S1006). (See below for further details.) Figure 9C The process is described in detail. When the determination in step S1005 is "No", CPU 201 skips the process in step S1006.
[0140] Subsequently, CPU 201 determines whether... Figure 1 The free mode switch has been operated on the first switch panel 102 (step S1007). When it is determined to be "yes", the CPU 201 executes free mode setting processing to change the value of the variable FreeMode on RAM 203 (step S1008). The free mode switch can be operated, for example, by toggle, and for example in Figure 9A In step S903, the initial value of the variable FreeMode is set to 1. When the free mode switch is pressed in this state, the value of the variable FreeMode becomes 0, and when the free mode switch is pressed again, the value of the variable FreeMode becomes 1. That is, the value of the variable FreeMode alternates between 0 and 1 each time the free mode switch is pressed. When the value of the variable FreeMode is 1, free mode is set; when the value is 0, free mode setting is canceled. When the determination in step S1007 is "No", CPU201 skips the processing of step S1008.
[0141] Subsequently, CPU 201 determines whether... Figure 1The performance rhythm adjustment switch has been operated on the first switch panel 102 (step S1009). When it is determined to be "yes", the CPU 201 executes the performance rhythm adjustment setting process, which changes the value of the variable ShiinAdjust on RAM 203 to the value specified by the numeric key on the first switch panel 102, followed by the operation of the performance rhythm adjustment switch (step S1010). For example, in Figure 9A In step S903, the initial value of the variable ShiinAdjust is set to 0. When the determination in step S1009 is "No", CPU 201 skips the processing of step S1010.
[0142] Finally, CPU 201 was determined. Figure 1 Checking whether any other switches on the first switch panel 102 or the second switch panel 103 have been operated, and performing the corresponding processing for each switch operation (step S1011). After this, CPU 201 terminates. Figure 8 The switching process in step S802, such as Figure 10 The flowchart is shown.
[0143] Figure 9B It is shown Figure 10 A flowchart illustrating a detailed example of the rhythm change processing in step S1002 is provided. As mentioned above, a change in the rhythm value also results in a change in TickTime (sec). Figure 9B In the flowchart shown, CPU 201 performs control processing related to changing TickTime (sec).
[0144] First, similar to in Figure 8 The initialization process executed in step S801 Figure 9A In step S901, CPU 201 calculates TickTime (sec) through arithmetic processing corresponding to equation (3) (step S911). Note that it is assumed that TickTime (sec) has already been used. Figure 1 The tempo value Tempo changed by the tempo change switch on the first switch panel 102 is stored in RAM 203, etc.
[0145] Next, similar to in Figure 8 The initialization process executed in step S801 Figure 9A In step S902, CPU 201 uses the TickTime (sec) calculated in step S911 to... Figure 2 Timer 210 in the CPU is set to interrupt the timer (step S912). Subsequently, CPU 201 terminates. Figure 10 The rhythm change processing in step S1002 is shown in the flowchart below. Figure 9B As shown.
[0146] Figure 9C It is shown Figure 10 The flowchart below provides a detailed example of step S1006, which is the start of song processing.
[0147] First, for automatic playback, CPU 201 initializes the values of timing data variables DeltaT_1 (first track block) and DeltaT_2 (second track block) on RAM 203 to count the relative time from the last event to 0, in units of TickTime. Next, CPU 201 initializes the corresponding values of variables AutoIndex_1 and AutoIndex_2 on RAM 203 to 0. Variable AutoIndex_1 is used to... Figure 7 The performance data in the first track block of the music clip data shown assigns the value of i to DeltaTime_1[i] and Event_1[i] (1≤i≤L-1), and the variable AutoIndex_2 is used to specify the value of i for DeltaTime_1[i] and Event_1[i], and the variable AutoIndex_2 is used to specify the value of i for the performance data in the first track block of the music clip data shown. Figure 7 The performance data in the second track block of the music segment data shown specifies the value of j (1≤j≤M-1) for DeltaTime_2[j] and Event_2[j] (as described in step S921). Therefore, in Figure 7 In the example, the performance data pairs DeltaTime_1[0] and Event_1[0] at the beginning of the first track block and the performance data pairs DeltaTime_2[0] and Event_2[0] at the beginning of the second track block are referred to as the initial states.
[0148] Next, CPU 201 initializes the value of the variable SongIndex on RAM 203, which specifies the current song position, to null (step S922). In many cases, null is usually defined as 0. However, since there is a case where the index number is 0, in this embodiment, null is defined as -1.
[0149] CPU 201 also initializes the value of the variable SongStart on RAM 203 to 1 (advance) (step S923), which indicates whether the lyrics and accompaniment are advanced (=1) or not advanced (=0).
[0150] Then, CPU 201 determines whether the user has already used [the system / mechanism]. Figure 1 The first switch panel 102 is set to reproduce the accompaniment in harmony with the playback of the lyrics (step S924).
[0151] When the determination in step S924 is "yes", CPU 201 sets the value of the variable Bansou in RAM 203 to 1 (accompaniment present) (step S925). Conversely, when the determination in step S924 is "no", CPU 201 sets the value of the variable Bansou to 0 (no accompaniment) (step S926). After processing in steps S925 or S926, CPU 201 terminates. Figure 10 The song processing begins in step S1006, as shown in the flowchart below. Figure 9C As shown.
[0152] Figure 11 It is shown Figure 8 A detailed flowchart of step S803's keyboard processing is provided. First, CPU 201 determines whether it has already been... Figure 2 The key scanner 206 in the middle operated Figure 1 Any key on the keyboard 101 (step S1101).
[0153] When the determination in step S1101 is "No", CPU 201 ends. Figure 8 The keyboard processing in step S803, such as Figure 11 The flowchart is shown in the image.
[0154] When the determination in step S1101 is "yes", the CPU 201 determines whether a key press operation or a key release operation has been performed (step S1102).
[0155] When it is determined in step S1102 that a key release operation has been performed, CPU 201 instructs speech synthesis LSI 205 to cancel the vocalization of the singing sound data 217 corresponding to the key release pitch (or key number) (step S1113). In response to this instruction, the speech synthesis LSI 205... Figure 3 The speech synthesis unit 302 stops emitting the corresponding singing sound data 217. Afterwards, the CPU 201 terminates... Figure 8 The keyboard processing in step S803, such as Figure 11 The flowchart is shown in the document.
[0156] When it is determined in step S1102 that a key press operation has been performed, CPU 201 determines the value of the variable FreeMode on RAM 203 (step S1103). The value of the variable FreeMode is as described above. Figure 10 The settings are configured in step S1008. When the value of the variable FreeMode is 1, free mode is set; when the value is 0, free mode is canceled.
[0157] When it is determined in step 1103 that the value of the variable FreeMode is 0 and the free mode setting has been cancelled, as mentioned above... Figure 6 As described in the performance style output unit 603, the CPU 201 sets the value calculated by the arithmetic processing shown in the following equation (4) using DeltaTime_1[AutoIndex_1], which will be described later, to the variable PlayTempo on RAM 203, indicating the corresponding... Figure 6 The performance style data 611 in A is the performance rhythm, and DeltaTime_1[AutoIndex_1] is each timing data 605 sequentially read from the music segment data 604 loaded into RAM 203 for song playback (step S1109).
[0158] [Expression 16]
[0159] PlayTempo=(1 / DeltaTime_1[AutoIndex_1])×Predetermined coefficient (4)
[0160] In equation (4), the predetermined coefficient in this embodiment is the TimeDivision value of the music segment data × 60. That is, if the TimeDivision value is 480, then when DeltaTime_1[AutoIndex_1] is 480, PlayTempo becomes 60 (corresponding to a normal tempo of 60). When DeltaTime_1[AutoIndex_1] is 240, PlayTempo becomes 120 (equivalent to a normal tempo of 120).
[0161] When the free mode setting is canceled, the performance rhythm is set to synchronize with timing information related to song playback.
[0162] When the value of the variable FreeMode is determined to be 1 in step 1103, the CPU 201 further determines whether the value of the variable NoteOnTime on RAM 203 is null (step S1104). This occurs at the start of song playback, for example, when... Figure 9A In step S903, the value of the variable NoteOnTime has been initially set to null, and after the song playback starts, it will be set sequentially in step S1110, which will be described later. Figure 2 The current time of timer 210 in the system.
[0163] When the song begins playing and when the determination in step S1104 is "yes", the playing rhythm cannot be determined based on the user's key press operation. Therefore, the CPU 201 sets the value calculated by the arithmetic processing shown in equation (4) using DeltaTime_1[AutoIndex_1], which is the timing data 605 on RAM 203, to the variable PlayTempo on RAM 203 (step S1109). In this way, when the song begins playing, the playing rhythm is temporarily set in a way that is synchronized with the timing information about the song playing.
[0164] After the song playback begins and when the determination in step S1104 is "No", the CPU 201 first sets the difference time to the variable DeltaTime on RAM 203, which is obtained by... Figure 2 The current time indicated by timer 210 is obtained by subtracting the value of the variable NoteOnTime in RAM 203, which indicates the time of the last key press (step S1105).
[0165] Next, CPU 201 determines whether the value of the variable DeltaTime, which indicates the difference time from the last key press time to the current key press time, is less than a predetermined maximum time for a key press to be considered as being played by a chord (step S1106).
[0166] When the determination in step S1106 is "yes" and it is determined that the current key press is a simultaneous key press during chord playing (chord), the CPU 201 does not perform the processing for determining the playing rhythm and proceeds to step S1110, which will be described later.
[0167] When the determination in step S1106 is "No" and it is determined that the current key press is not a simultaneous key press during chord playing (chord), the CPU 201 further determines whether the value of the variable DeltaTime, which indicates the difference time from the last key press to the current key press, is greater than the minimum time used to consider the performance as interrupted (step S1107).
[0168] If step S1107 determines "yes" and it is determined that the key press occurred after a period of interruption in the performance (the beginning of a musical phrase), then the rhythm of the musical phrase cannot be determined. Therefore, CPU 201 sets the value calculated using DeltaTime_1[AutoIndex_1], which is the timing data on RAM 203, through the arithmetic processing shown in equation (4) to the variable PlayTempo on RAM 203 (step S1109). In this way, in the case of a key press occurring after a period of interruption in the performance (the beginning of a musical phrase), the performance rhythm is temporarily set in a way that is synchronized with the timing information related to the song playback.
[0169] When the determination in step S1107 is "No" and it is determined that the current key press is neither providing chord playing (chord) nor is it a key press at the beginning of a musical phrase, the CPU 201 sets the value obtained by multiplying a predetermined coefficient by the reciprocal of the variable DeltaTime, which indicates the difference time from the last key press to the current key press (as shown in equation (5) below), to the indicator on RAM 203 corresponding to... Figure 6 The performance style data 611 in the performance time is used to determine the variable PlayTempo of the performance rhythm (step S1108).
[0170] [Expression 17]
[0171] PlayTempo=(1 / DeltaTime)×Pre-booking coefficient (5)
[0172] As a result of the processing in step S1108, when the value of the variable DeltaTime, which indicates the time difference between the last key press and the current key press, is small, the value of PlayTempo, which represents the playing rhythm, increases (the playing rhythm becomes faster), the playing phrase is considered a fast segment, and in the speech synthesis unit 302 of the speech synthesis LSI 205, the sound waveform of the singing voice data 217 is inferred, wherein the duration of the consonant part is shorter, such as Figure 5A As shown. On the other hand, when the value of the variable DeltaTime, which indicates the difference in time, is large, the value of the performance rhythm becomes small (the performance rhythm becomes slow), and the performance phrase is regarded as a slow passage. In speech synthesis, in part 302, the sound waveform of the singing voice data 217 is inferred, where, as Figure 5B As shown, the duration of the consonant part is relatively long.
[0173] After the processing in step S1108, after the processing in step S1109, or after the determination in step S1106 becomes "yes", the CPU 201 will be... Figure 2The current time indicated by timer 210 in the RAM is set to the variable NoteOnTime on RAM 203, which indicates the time of the last key press (step S1110).
[0174] Finally, CPU 201 will add the value of the variable ShiinAdjust on RAM 203 (see [link]). Figure 10 The value obtained in step S1010 is set as a new value for the variable PlayTempo (step S1111), wherein the performance rhythm adjustment value intentionally set by the user is set to the value of the variable PlayTempo on RAM 203, which indicates the performance rhythm determined in step S1108 or S1109. After this, CPU 201 terminates. Figure 8 The keyboard processing in step S803, such as Figure 11 As shown in the flowchart.
[0175] Through the processing in step S1111, the user can intentionally adjust the duration of the consonant portions in the singing voice data 217 synthesized in the speech synthesis unit 302. In some cases, the user may want to adjust the singing style, depending on the song title or taste. For example, for some songs, when the user wants to provide a performance with good vocal quality by shortening the overall sound, the user may want to generate a speech sound by shortening the consonants, as if singing the song in a fast speaking manner. Conversely, for some songs, when the user wants to perform smoothly overall, the user may want to generate a speech sound that clearly conveys the breath of the consonants, as if singing a song slowly. Therefore, in this embodiment, the user can, for example, operate... Figure 1 The playing tempo adjustment switch on the first switch panel 102 changes the value of the variable ShiinAdjust, and based on this, the value of the variable PlayTempo is adjusted to synthesize singing sound data 217 that reflects the user's intention. In addition to the switch operation, the value of ShiinAdjust can be precisely controlled at any time during a piece of music by operating a pedal with a variable resistor connected to the electronic keyboard instrument 100 with the foot.
[0176] The tempo value set for the variable PlayTempo via the keyboard processing described above is set as part of the vocal data 215 during the song playback processing described later (see later description). Figure 13 The process is carried out in step S1305 and then published to the speech synthesis LSI 205.
[0177] In the above keyboard processing, in particular, the processing of steps S1103 to S1109 and step S1111 corresponds to Figure 6 The function of the performance style output unit 603 in the middle.
[0178] Figure 12 This shows the time based on each TickTime (sec) by Figure 2 A detailed flowchart illustrating the automatic performance interrupt handling process executed due to an interrupt generated by Timer 210 (see [reference]). Figure 9A Step S902 or Figure 9B Step S912). Figure 7 The performance data of the first and second track blocks in the music segment data shown are processed as follows.
[0179] First, CPU 201 executes a series of processes corresponding to the first track block (steps S1201 to S1206). First, CPU 201 determines whether the value of SongStart is 1 (see...). Figure 10 Step S1006 and Figure 9C In step S923, that is, whether the lyrics and accompaniment have been indicated (step S1201).
[0180] When it is determined that there is no indication of lyrics and accompaniment (the determination in step S1201 is "No"), CPU 201 ends. Figure 12 The flowchart shows the automatic performance interruption process, without the lyrics and accompaniment being played.
[0181] When it is determined that the lyrics and accompaniment have been instructed (determined as "yes" in step S1201), CPU 201 determines whether the value of variable DeltaT_1 on RAM 203 matches DeltaTime_1[AutoIndex_1] on RAM 203 (step S1202), where variable DeltaT_1 indicates the relative time relative to the first track block since the last event, and DeltaTime_1[AutoIndex_1] is timing data 605 ( Figure 6 ), which indicates the waiting time for the performance data pair to be executed, represented by the value of the variable AutoIndex_1 on RAM 203.
[0182] When the determination in step S1202 is "No", CPU 201 increments the value of variable DeltaT_1 by 1. This variable DeltaT_1 indicates the relative time since the last event relative to the first track block, and allows the time to be advanced to correspond to 1 TickTime unit of the current interrupt (step S1203). Afterward, CPU 201 proceeds to step S1207, which will be described later.
[0183] When the determination in step S1202 is "yes", CPU 201 stores the value of variable AutoIndex_1 in variable SongIndex on RAM 203 (step S1204). Variable AutoIndex_1 indicates the position of the song event that should be executed next in the first track block.
[0184] In addition, CPU 201 increments the value of the variable AutoIndex_1, which is used to reference the performance data pairs in the first track block, by 1 (step S1205).
[0185] Furthermore, CPU 201 resets the value of variable DeltaT_1 to 0 (step S1206), which indicates the relative time since the song event most recently referenced in the first track block. Afterward, CPU 201 proceeds to step S1207.
[0186] Next, CPU 201 executes a series of processes corresponding to the second track block (steps S1207 to S1213). First, CPU 201 determines whether the value of the variable DeltaT_2 on RAM 203 matches DeltaTime_2[AutoIndex_2] on RAM 203. The variable DeltaT_2 indicates the relative time to the second track block since the last event. DeltaTime_2[AutoIndex_2] is the timing data of the performance data pair to be executed, indicated by the value of the variable AutoIndex_2 on RAM 203 (step S1207).
[0187] When the determination in step S1207 is "No", CPU 201 increments the value of variable DeltaT_2 by 1. DeltaT_2 indicates the relative time since the last event relative to the second track block, allowing time to be advanced by one TickTime unit corresponding to the current interrupt (step S1208). After this, CPU 201 terminates. Figure 12 The flowchart shows the automatic performance interruption handling.
[0188] When the determination in step S1207 is "yes", CPU 201 determines whether the value of the variable Bansou on RAM 203, which indicates accompaniment playback, is 1 (accompaniment present) or not 1 (accompaniment absent) (step S1209) (see step S1209). Figure 9C Steps S924 to S926 in the process.
[0189] When the determination in step S1209 is "Yes", CPU 201 executes the processing indicated by event data Event_2[AutoIndex_2] on RAM 203 related to the accompaniment of the second track block indicated by the value of variable AutoIndex_2 (step S1210). When the processing indicated by event data Event_2[AutoIndex_2] executed here is, for example, a note-on event, the key number and velocity specified by the note-on event are used to... Figure 2 The LSI 204, the sound source in the instrument, issues instructions to generate musical sounds for accompaniment. On the other hand, when the process indicated by the event data Event_2[AutoIndex_2] is, for example, a note-off event, the key number specified by the note-off event is used to... Figure 2 The LSI 204 sound source in the middle issues a command to cancel the music sound of the accompaniment that is being generated.
[0190] On the other hand, when the determination in step S1209 is "no", the CPU 201 skips step S1210 and proceeds to the next step S1211 to process in sync with the lyrics without executing the processing indicated by the event data Event_2[AutoIndex_2] related to the current accompaniment, and only executes the control processing of the advance event.
[0191] After step S1210, or when the determination in step S1209 is "No", the CPU 201 increments the value of the variable AutoIndex_2, which is used to reference the performance data pair on the second track block, by 1 (step S1211).
[0192] Next, CPU 201 resets the value of the variable DeltaT_2, which indicates the relative time since the most recent event executed for the second track block, to 0 (step S1212).
[0193] Then, CPU 201 determines whether the value of timing data DeltaTime_2[AutoIndex_2] on RAM 203 for the performance data pair on the second track block to be executed next, as indicated by the value of variable AutoIndex_2, is 0, that is, whether the event will be executed at the same time as the current event (step S1213).
[0194] When the determination in step S1213 is "No", CPU 201 ends. Figure 12 The flowchart in the image shows the current automatic performance interruption process.
[0195] When the determination in step S1213 is "yes", CPU 201 returns to the processing of step S1209 and repeats the control processing related to the event data Event_2[AutoIndex_2] in the performance data pair to be executed next on RAM 203 for the second track block indicated by the value of variable AutoIndex_2. CPU 201 repeats the processing of steps S1209 to S1213 up to the number of times it needs to be executed simultaneously. When multiple note activation events are about to produce sound at the same time, such as chords, the above processing sequence is executed.
[0196] Figure 13 It is shown Figure 8 The flowchart shows a detailed example of the song playback processing in step S805.
[0197] First of all, Figure 12 In step S1204 of the automatic performance interruption handling, CPU 201 determines whether a new value other than null has been set for the variable SongIndex on RAM 203 to enter the song playback state (step S1301). For the variable SongIndex, initially at the start of the song... Figure 9C In step S922, an empty value is set. In step S1204, a valid value is set for the variable AutoIndex_1, which indicates the position of the next song event to be executed in the first track block. This value is set whenever the singing sound playback timer reaches the specified value. Figure 12 In the automatic performance interruption handling, if the determination in step S1202 is "yes", continue; each time it is executed again... Figure 13 In the song playback process shown in the flowchart, a null value is set again in step S1307, which will be described later. That is, whether the value of the variable SongIndex is set to a valid non-null value indicates whether the current timing is for song playback.
[0198] When the determination in step S1301 is "yes", that is, when the current time is the song playback timer, CPU 201 determines... Figure 8 The keyboard processing in step S803 has been detected. Figure 1 The new user key on keyboard 101 is pressed (step S1302).
[0199] When the determination in step S1302 is "yes", the CPU 201 will set the specified pitch by pressing the user key to a variable in a register or RAM 203 (not specifically shown) as the sound pitch (step S1303).
[0200] On the other hand, when it is determined through the determination in step S1301 that the current time is the song playback timing and the determination in step S1302 is "no", that is, it is determined that no new key press has been detected at the current time, the CPU 201 reads the pitch data (corresponding to) from the song event data Event_1[SongIndex] on the first track block of music segment data in RAM 203, indicated by the variable SongIndex in RAM 203. Figure 6 The event data 606 contains pitch data 607, and the pitch data is set to a variable in a register or RAM 203 that is not specifically shown (step S1304).
[0201] Subsequently, CPU 201 reads the lyrics string (corresponding to the song event Event_1[SongIndex] on the first track block of music segment data in RAM 203, as indicated by the variable SongIndex on RAM 203) Figure 6 The event data 606 contains the lyrics data 608. Then, the CPU 201 sets the singing voice data 215 during performance, which reads the lyrics string (corresponding to...). Figure 6 The lyrics data 609 during performance), and the pitch of the sound obtained in step S1303 or S1304 (corresponding to the performance data 609). Figure 6 The pitch data during performance (610), and in the corresponding Figure 8 Step S803 in Figure 10 The playing rhythm obtained in step S1111 for the variable PlayTempo on RAM 203 (corresponding to) Figure 6 The performance style data (611) is set to a variable in a register or RAM 203 that is not specifically shown (step S1305).
[0202] Subsequently, CPU 201 publishes the singing voice data 215 generated in step S1305 during the performance to... Figure 2 LSI 205 speech synthesis Figure 3 The speech synthesis unit 302 (step S1306) is mentioned. See reference... Figures 3 to 6 As described, the speech synthesis LSI 205 infers, synthesizes, and outputs singing voice data 217 based on lyrics specified by singing voice data 215 during performance. The singing voice data 217 corresponds in real-time to pitch data 607 automatically assigned by user key presses on keyboard 101 or song playback as specified by the singing voice data 215 during performance (see reference). Figure 6 The pitch of the song is determined by the singing voice data 217, and the singing voice data 217 is used to sing the song appropriately according to the performance rhythm (singing style) specified by the singing voice data 215 during performance.
[0203] Finally, CPU 201 clears the value of the variable SongIndex to make it null and sets the subsequent timing to a non-song playback timing (step S1307). After this, CPU 201 terminates. Figure 8 The song playback processing in step S805, such as Figure 13 The flowchart is shown.
[0204] In the above song playback process, specifically, steps S1302 to S1304 correspond to... Figure 6 The function of the pitch designation unit 602 in the text. Specifically, the processing in step S1305 corresponds to... Figure 6 The function of the lyrics output unit 601 in the text.
[0205] According to the above embodiments, depending on the type of musical segment to be played and the musical phrase to be played, for example, the duration of the consonant part in the human voice is longer in a slow passage with fewer notes, which can produce a highly expressive and lively sound, and shorter in a fast-paced passage or with many notes, which can produce a clear sound. That is, a change in timbre that matches the musical phrase to be played can be obtained.
[0206] The above embodiments are examples of electronic musical instruments configured to generate singing voice data. However, as another embodiment, embodiments of electronic musical instruments configured to generate the sounds of wind or string instruments can also be implemented. In this case, corresponding to Figure 3 The acoustic model unit 306 stores a trained acoustic model that performs machine learning using training pitch data for a specified pitch, teacher data corresponding to training acoustic data indicating the acoustics of a specific sound source for a wind or string instrument corresponding to the pitch, and training performance style data representing the performance style (e.g., performance rhythm) of the training acoustic data. It outputs acoustic model parameters corresponding to the input pitch data and performance style data. Furthermore, the pitch specification unit (corresponding to...) Figure 6 The pitch specification unit 602 outputs performance pitch data indicating the pitch specified by the user's performance operation during performance. Further, the performance style output unit (corresponding to...) Figure 6 The performance style output unit 603 in the middle outputs performance style data indicating the performance style during the performance, such as performance rhythm. The sound generation model unit (corresponding to...) Figure 3The acoustic model unit 308 synthesizes and outputs musical sound data. This musical sound data is based on acoustic model parameters output from a trained acoustic model stored in the acoustic model unit, obtained by inputting pitch data and performance style data during performance. During performance, the speech sound of a sound source is inferred. In this embodiment of the electronic instrument, for example, in a song with a fast passage, pitch data such as the sound of a wind instrument or the speed at which a bow strikes a stringed instrument as if the bow is being struck at a slower speed makes a clear performance possible. Conversely, in a song with a low passage, pitch data such as the sound of a wind instrument or the lengthened time of a bow strike makes a highly expressive performance possible.
[0207] In the above embodiments, when the speed of a musical phrase cannot be estimated, such as the first key press or the first key press of a musical phrase, the rising part of a consonant or sound becomes shorter when the singing or striking sound is strong, and longer when the singing or striking sound is weak. Utilizing this trend, the force of playing the keyboard (the speed at which the key is pressed) can be used as a basis for calculating the rhythm value of the performance.
[0208] It can be used as Figure 3 The speech synthesis method of the vocalization model unit 308 is not limited to cepstral speech synthesis method, but can adopt a variety of speech synthesis methods including LSP speech synthesis method.
[0209] In addition, as a speech synthesis method, any speech synthesis method can be used, except for those based on statistical speech synthesis processing using HMM acoustic models and those based on statistical speech synthesis processing using DNN acoustic models, as long as it is a technique that uses machine learning-based statistical speech synthesis processing, such as combining HMM and DNN acoustic models.
[0210] In the above embodiment, the lyrics data 609 is provided as pre-stored music segment data 604 during performance. However, text data obtained by performing speech recognition on the user's real-time singing can be provided as real-time lyrics information.
[0211] In connection with the above embodiments, the following appendix is also disclosed.
[0212] (Appendix 1)
[0213] An electronic musical instrument, comprising:
[0214] A pitch specification unit, which is configured to output performance pitch data specified during performance;
[0215] The performance style output unit is configured to output performance style data indicating the performance style during the performance; and
[0216] The sound generation model unit is configured to synthesize and output musical sound data corresponding to the pitch data and performance style data at the time of performance, based on acoustic model parameters inferred by inputting the pitch data and performance style data at the time of performance into a trained acoustic model.
[0217] (Appendix II)
[0218] An electronic musical instrument, comprising:
[0219] The lyrics output unit is configured to output performance-time lyrics data that indicates the lyrics during performance.
[0220] A pitch specification unit, which is configured to output performance pitch data that is tuned to the output of lyrics during performance;
[0221] The performance style output unit is configured to output performance style data indicating the performance style during the performance; and
[0222] A vocal model unit is configured to synthesize and output singing voice data corresponding to the lyrics data, pitch data, and performance style data at the time of performance during performance, based on acoustic model parameters inferred by inputting the lyrics data, pitch data, and performance style data at the time of performance into a trained acoustic model.
[0223] (Appendix 3)
[0224] According to the electronic musical instrument described in Appendix 1 or 2, the performance style output unit is configured to sequentially measure the time intervals of specified pitches during performance and sequentially output performance rhythm data indicating the sequentially measured time intervals as performance style data during performance.
[0225] (Appendix 4)
[0226] According to the electronic musical instrument described in Appendix 3, the performance style output unit includes a changing device for allowing a user to intentionally change the sequentially acquired performance rhythm data.
[0227] (Appendix 5)
[0228] An electronic musical instrument control method includes causing the electronic musical instrument's processor to perform the following processes:
[0229] Output the pitch data specified during the performance;
[0230] Output performance style data indicating the performance style during the performance; and
[0231] Based on acoustic model parameters inferred by inputting the pitch data and performance style data during the performance into a trained acoustic model, musical sound data corresponding to the pitch data and performance style data during the performance is synthesized and output during the performance.
[0232] (Appendix 6)
[0233] An electronic musical instrument control method includes causing the electronic musical instrument's processor to perform the following processes:
[0234] Output the lyrics data during the performance, indicating the lyrics to be played.
[0235] Output the specified pitch data for the performance, tuned to the output of the lyrics during the performance;
[0236] Output performance style data indicating the performance style during the performance; and
[0237] Based on acoustic model parameters inferred by inputting the lyrics data, pitch data, and performance style data during the performance into a trained acoustic model, singing voice data corresponding to the lyrics data, pitch data, and performance style data during the performance are synthesized and output during the performance.
[0238] (Appendix 7)
[0239] A program that enables the processor of an electronic musical instrument to perform the following processes:
[0240] Output the pitch data specified during the performance;
[0241] Output performance style data indicating the performance style during the performance; and
[0242] Based on acoustic model parameters inferred by inputting the pitch data and performance style data during the performance into a trained acoustic model, musical sound data corresponding to the pitch data and performance style data during the performance is synthesized and output during the performance.
[0243] (Appendix 8)
[0244] A program that enables the processor of an electronic musical instrument to perform the following processes:
[0245] Output the lyrics data during the performance, indicating the lyrics to be played.
[0246] Output the specified pitch data for the performance, tuned to the output of the lyrics during the performance;
[0247] Output performance style data indicating the performance style during the performance; and
[0248] Based on acoustic model parameters inferred by inputting the lyrics data, pitch data, and performance style data during the performance into a trained acoustic model, singing voice data corresponding to the lyrics data, pitch data, and performance style data during the performance are synthesized and output during the performance.
[0249] This application is based on Japanese Patent Application No. 2020-152926, filed on September 11, 2020, the contents of which are incorporated herein by reference.
[0250] Reference Symbol List
[0251] 100: Electronic keyboard instruments
[0252] 101: Keyboard
[0253] 102: First switch panel
[0254] 103: Second switch panel
[0255] 104: LCD
[0256] 200: Control System
[0257] 201: CPU
[0258] 202: ROM
[0259] 203: RAM
[0260] 204: Sound Source LSI
[0261] 205: Sound Synthesis LSI
[0262] 206: Key Scanner
[0263] 208: LCD Controller
[0264] 209: System Bus
[0265] 210: Timer
[0266] 211, 211: D / A converter
[0267] 213: Mixer
[0268] 214: Amplifier
[0269] 215: Singing Voice Data
[0270] 216: Sound Generation Control Data
[0271] 217: Singing Voice Data
[0272] 218: Music Sound Data
[0273] 219: Network Interface
[0274] 300: Server computer
[0275] 301: Voice Training Department
[0276] 302: Sound Synthesis Department
[0277] 303 Training Singing Voice Analysis Unit
[0278] 304: Training Acoustic Feature Extraction Unit
[0279] 305: Model Training Unit
[0280] 306: Acoustic Model Unit
[0281] 307: Vocal Analysis Unit During Performance
[0282] 308: Sound-generating model unit
[0283] 309: Sound Source Generation Unit
[0284] 310: Synthetic Filter Unit
[0285] 311: Training Singing Voice Data
[0286] 312: Training Singing Voice Data
[0287] 313: Training Language Feature Sequences
[0288] 314: Training Acoustic Feature Sequences
[0289] 315: Training Result Data
[0290] 316: Language Feature Sequence During Performance
[0291] 317: Acoustic characteristic sequence during performance
[0292] 318: Spectrum Information
[0293] 319: Sound Source Information
[0294] 601: Lyrics Output Unit
[0295] 602: Designated unit of pitch
[0296] 603: Performance Style Output Unit
[0297] 604: Music clip data
[0298] 605: Timed Data
[0299] 606: Event Data
[0300] 607: Pitch Data
[0301] 608: Lyrics Data
[0302] 609: Lyrics data during performance
[0303] 610: Pitch data during performance
[0304] 611: Performance style data during performance
Claims
1. An electronic musical instrument, comprising: The pitch designation unit is configured to output pitch data during performance in response to the user's performance operation. A performance style output unit is configured to generate and output performance style data indicating the performance speed by sequentially and in real-time measuring the time intervals between consecutive operations performed by the user during the performance, and by calculating the performance speed based on the sequentially measured time intervals; and A sound generation model unit is configured to synthesize and output musical sound data corresponding to the pitch data and performance style data at the time of performance, based on acoustic model parameters inferred by inputting the pitch data and performance style data at the time of performance into a trained acoustic model.
2. An electronic musical instrument, comprising: The lyrics output unit is configured to output performance-time lyrics data that indicates the lyrics during performance. A pitch designation unit is configured to output performance pitch data that is tuned to the output of lyrics in response to a user's performance operation during the performance. A performance style output unit is configured to generate and output performance style data indicating the performance speed by sequentially and in real-time measuring the time intervals between consecutive operations performed by the user during the performance, and by calculating the performance speed based on the sequentially measured time intervals; and A vocal model unit is configured to synthesize and output singing voice data corresponding to the lyrics data, pitch data, and performance style data during the performance, based on acoustic model parameters inferred by inputting the lyrics data, pitch data, and performance style data during the performance into a trained acoustic model.
3. The electronic musical instrument of claim 1 or 2, wherein, The performance style output unit includes a changing device for allowing a user to intentionally change the sequentially acquired performance rhythm data.
4. A method for controlling an electronic musical instrument, comprising causing the processor of the electronic musical instrument to perform the following processes: In response to the user's playing actions during performance, output the pitch data during performance; By sequentially measuring the time intervals between consecutive operations performed by the user during the performance in real time, and calculating the performance tempo based on the sequentially measured time intervals, performance style data indicating the performance tempo is generated and output; and Based on acoustic model parameters inferred by inputting the pitch data and performance style data during the performance into a trained acoustic model, musical sound data corresponding to the pitch data and performance style data during the performance is synthesized and output during the performance.
5. A method for controlling an electronic musical instrument, comprising causing the processor of the electronic musical instrument to perform the following processes: Output the lyrics data during performance, indicating the lyrics at the time of performance; In response to the user's performance operation during the performance, output performance pitch data that is tuned to the output of lyrics; By sequentially measuring the time intervals between consecutive operations performed by the user during the performance in real time, and calculating the performance tempo based on the sequentially measured time intervals, performance style data indicating the performance tempo is generated and output; and Based on acoustic model parameters inferred by inputting the lyrics data, pitch data, and performance style data during the performance into a trained acoustic model, singing voice data corresponding to the lyrics data, pitch data, and performance style data during the performance are synthesized and output during the performance.
6. A non-transitory computer-readable storage medium storing a program that causes a processor of an electronic musical instrument to perform the following processes: In response to the user's playing actions during performance, output the pitch data during performance; By sequentially measuring the time intervals between consecutive operations performed by the user during the performance in real time, and calculating the performance tempo based on the sequentially measured time intervals, performance style data indicating the performance tempo is generated and output; and Based on acoustic model parameters inferred by inputting the pitch data and performance style data during the performance into a trained acoustic model, musical sound data corresponding to the pitch data and performance style data during the performance is synthesized and output during the performance.
7. A non-transitory computer-readable storage medium storing a program that causes a processor of an electronic musical instrument to perform the following processes: Output the lyrics data during performance, indicating the lyrics at the time of performance; In response to the user's performance operation during the performance, output performance pitch data that is tuned to the output of lyrics; By sequentially measuring the time intervals between consecutive operations performed by the user during the performance in real time, and calculating the performance tempo based on the sequentially measured time intervals, performance style data indicating the performance tempo is generated and output; and Based on acoustic model parameters inferred by inputting the lyrics data, pitch data, and performance style data during the performance into a trained acoustic model, singing voice data corresponding to the lyrics data, pitch data, and performance style data during the performance are synthesized and output during the performance.