Electronic musical instrument, control method of electronic musical instrument, and storage medium

By combining multiple performance operators and machine learning acoustic models in electronic musical instruments, musical sound data that simulates professional performers is generated, solving the problem of beginners having difficulty reproducing pitch, duration, and beat accents when playing music, and improving the performance effect.

CN113874932BActive Publication Date: 2025-09-23CASIO COMPUTER CO LTD
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202080038164.7
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Priority Date
2019-05-23
Filing Date
2020-05-15
Publication Date
2025-09-23
Estimated Expiration
2040-05-15

AI Technical Summary

Technical Problem

It is difficult for beginners to reproduce the professional performers' exquisite performance of the pitch, duration, beat emphasis, etc. of each note in the phrase when playing music.

Method used

Electronic musical instruments are associated with pitch data through multiple performance operators, combined with memory and processors, and use acoustic models trained by machine learning to generate inferred musical data based on user operations to simulate the performance skills of professional performers.

Benefits of technology

It can produce the same sound as that played by professional players based on simple user operations, improving the performance of beginners.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN113874932B_ABST
    Figure CN113874932B_ABST
Patent Text Reader

Abstract

An electronic musical instrument according to an example of an embodiment inputs pitch data (215) corresponding to a user operation on a certain performance operator among a plurality of performance operators (101) into a learned acoustic model (606), outputs acoustic feature quantity data (617) corresponding to the pitch data (215) from the learned acoustic model (606), and digitally synthesizes and outputs inference music data (217) including an inference performance technique of a certain musician, the inference music data being based on the output acoustic feature quantity data (617) and not executed by user operation.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to an electronic musical instrument configured to reproduce musical instrument sounds in response to an operation of an operator such as a keyboard, a control method of the electronic musical instrument, and a storage medium. Background Art

[0002] With the widespread use of electronic musical instruments, users can now enjoy a variety of musical performances. For example, even beginners can easily enjoy playing a piece of music by following and touching the illuminated keys or by following the performance instructions displayed on the display. The pitch, duration, note-on timing, and beat emphasis of the notes played by the user depend on the user's performance skills.

[0003] Citation List

[0004] Patent Literature

[0005] Patent Document 1: JP-H09-050287A Summary of the Invention

[0006] Technical issues

[0007] Especially when a user is a beginner, when playing phrases of various pieces of music on a musical instrument, the user tries to follow the pitch of each note in the musical score. However, it is difficult to reproduce the instrument-appropriate and exquisite performance of a professional player in terms of note-on timing, duration, beat emphasis, etc. of each note in the phrase.

[0008] Problem Solution

[0009] An electronic musical instrument according to one aspect of the present invention comprises:

[0010] a plurality of performance operators, each performance operator being associated with different pitch data;

[0011] A memory storing a trained acoustic model obtained by performing machine learning on:

[0012] A training music score dataset, including training pitch data; and

[0013] A training performance dataset obtained from performers playing their instruments; and

[0014] at least one processor,

[0015] wherein the at least one processor is configured to:

[0016] inputting pitch data corresponding to the user operation on one of the performance operators into the trained acoustic model according to a user operation on the one of the performance operators, so that the trained acoustic model outputs acoustic feature data corresponding to the input pitch data, and

[0017] Digitally synthesize and output inference sound data, which includes the performer's inference performance skills

[0018] The inferred musical tone data is based on acoustic feature data output by a trained acoustic model according to input pitch data, and

[0019] The inference musical tone data is not played during the user operation.

[0020] Advantageous Effects of the Invention

[0021] By implementing the present invention, it is possible to provide an electronic musical instrument configured to produce sounds as if played by a professional player according to user operations on performance operators. BRIEF DESCRIPTION OF THE DRAWINGS

[0022] Figure 1 An example of the appearance of an embodiment of an electronic keyboard instrument is shown.

[0023] Figure 2 is a block diagram showing a hardware configuration example of an embodiment of a control system for an electronic keyboard musical instrument.

[0024] Figure 3 is a block diagram showing a configuration example of a circulator LSI.

[0025] Figure 4 is a timing chart of the first embodiment of the loop recording / playback process.

[0026] Figure 5 The quantization process is illustrated.

[0027] Figure 6 is a block diagram showing a configuration example of a voice training unit and a voice synthesis unit.

[0028] Figure 7 A first embodiment of a statistical sound synthesis process is illustrated.

[0029] Figure 8 A second embodiment of the statistical sound synthesis process is illustrated.

[0030] Figure 9 1 is a main flowchart showing an example of control processing of an electronic musical instrument according to the second embodiment of the loop recording / playback processing.

[0031] Figure 10Ais a flowchart showing a detailed example of the initialization process.

[0032] Figure 10B is a flowchart showing a detailed example of the tempo change process.

[0033] Figure 11 is a flowchart showing a detailed example of the switching process.

[0034] Figure 12 is a flowchart showing a detailed example of tick time interrupt processing.

[0035] Figure 13 is a flowchart showing a detailed example of the pedal control process.

[0036] Figure 14 is a first flowchart showing a detailed example of the circulator control process.

[0037] Figure 15 is a second flowchart showing a detailed example of the circulator control process.

[0038] Figure 16 is a first timing chart of the second embodiment of the loop recording / playback process.

[0039] Figure 17 is a second timing chart of the second embodiment of the loop recording / playback process. DETAILED DESCRIPTION

[0040] Hereinafter, embodiments of the present invention will be described in detail with reference to the accompanying drawings.

[0041] Figure 1 An example of the appearance of an embodiment of an electronic keyboard instrument 100 is shown. The electronic keyboard instrument 100 includes: a keyboard 101 comprising a plurality of keys as a performance operator (plural performance operators 101); a first switch panel 102 for various settings, such as volume control and rhythm settings for loop recording; a second switch panel 103 for selecting the timbre of the sound module of the electronic keyboard instrument 100, the instrument to be synthesized, etc.; and a liquid crystal display (LCD) 104 configured to display various setting data, etc. A foot pedal 105 (pedal operator 105) for loop recording and playback is connected to the electronic keyboard instrument 100 via a cable. Furthermore, although not specifically shown, the electronic keyboard instrument 100 includes a speaker configured to produce musical sounds generated by the performance provided on the back surface, side surface, or rear surface, etc.

[0042] Figure 2 Shown Figure 1 This is a hardware configuration example of an embodiment of the control system 200 of the electronic keyboard instrument 100. Figure 2In the control system 200, connected to the system bus 209 are: a central processing unit (CPU) 201; a ROM (read-only memory / large-capacity flash memory) 202; a RAM (random access memory) 203; a sound module large-scale integration (LSI) 204; a sound synthesis LSI 205 (processor 205); a circulator LSI 220; a key scanner 206, Figure 1 The keyboard 101, the first switch panel 102, the second switch panel 103 and the pedal 105 are connected to the key scanner 206; and the LCD controller 208, Figure 1 The LCD 104 is connected to the LCD controller 208. In addition, a timer 210 for controlling loop recording / playback processing in the looper LSI 220 is connected to the CPU 201. The tone output data 218 output from the sound module LSI 204 and the loop playback inference tone data 222 output from the looper LSI 220 are mixed in the mixer 213 and then converted into an analog output signal by the digital-to-analog converter 211. The analog output signal is amplified in the amplifier 214 and then output from a speaker (not shown) or an output terminal (not shown).

[0043] The CPU 201 is configured to execute a control program stored in the ROM 202 by using the RAM 203 as a working memory. Figure 1 The ROM 202 stores control programs and various fixed data, as well as model parameters and the like as a result of training by machine learning described later.

[0044] The timer 210 used in the present embodiment is implemented on the CPU 201 and is configured to manage the progress of loop recording / playback in the electronic keyboard instrument 100 , for example.

[0045] The sound module LSI 204 is configured to load the tone output data 218 from a waveform ROM (not shown), for example, and output it to the mixer 213 according to the sound generation control data from the CPU 201. The sound module LSI 204 can generate up to 256 sounds simultaneously.

[0046] The sound synthesis LSI 205 receives data representing musical instruments and pitch data 215, which is a sequence of pitches for each musical phrase, from the CPU 201 in advance. The sound synthesis LSI 205 then digitally synthesizes inference music data 217 for the musical phrase, including performance expression sounds (the performer's inference performance techniques). The performance expression sounds represent sounds corresponding to performance techniques not performed by the user, such as the articulation of legato, a notation used in Western musical notation. The sound synthesis LSI 205 is configured to output the digitally synthesized inference music data 217 to the looper LSI 220.

[0047] Response pair Figure 1 The looper LSI 220 loop-records the inference musical tone data 217 output by the sound synthesis LSI 205 and the loop playback sound that is repeatedly played back in accordance with the operation of the pedal 105. The looper LSI 220 repeatedly outputs the loop playback inference musical tone data 222 obtained in the end to the mixer 213.

[0048] To communicate state changes by interrupting the CPU 201, the key scanner 206 is configured to scan: Figure 1 a key pressing / releasing state of the keyboard 101; a switch state of the first switch panel 102 and the second switch panel 103; and a pedal state of the pedal 105.

[0049] The LCD controller 208 is an integrated circuit (IC) configured to control the display state of the LCD 104 .

[0050] Figure 3 It shows Figure 2 The circulator LSI 220 is a block diagram of a configuration example of the circulator LSI 220. The circulator LSI 220 includes: a first loop storage area 301 and a second loop storage area 302 for storing loop playback inference musical tone data 222 of a repeated portion to be repeatedly played back; a loop recording unit 303 configured to record the inference musical tone data 217 output from the sound synthesis LSI 205 on one of the loop storage areas through a mixer 307; a loop playback unit 304 configured to play back the inference musical tone data 217 stored in one of the loop storage areas as a loop playback sound 310; a phrase delay unit 305 configured to delay the loop playback sound 310 by one phrase (one measure); and a mixer 307 configured to combine the loop playback sound delayed output 311 output from the phrase delay unit 305 with the loop playback sound delayed output 311 output from the sound synthesis LSI 205, and the inferred musical sound data 217 input is mixed and the mixed data is output to the loop recording unit 303; and the beat extraction unit 306 is configured to extract the beat timing from the loop playback sound 310 output by the loop playback unit 304 as beat data 221 and output it to the sound synthesis LSI 205.

[0051] In this embodiment, for example, Figure 1In the electronic keyboard instrument 100 of the present invention, as the automatic rhythm or accompaniment sound is generated according to known techniques and output from the sound module LSI 204, the user sequentially touches the keys on the keyboard 101 corresponding to the pitch sequence of a phrase in a musical score. The keys to be touched by the user can be guided by the illuminated keys on the keyboard 101 of known techniques. The user does not need to follow the pressing timing or pressing duration of the note keys. The user only needs to touch the keys at least following the pitch, that is, for example, by touching the keys only following the keys with light. Whenever the user completes pressing the keys in the phrase, Figure 2 The CPU 201 outputs the pitch sequence of the key combination in the phrase obtained by detecting the depression of the key to the sound synthesis LSI 205. As a result, after delaying one phrase, the sound synthesis LSI 205 can generate inference tone data 217 of the phrase based on the pitch sequence of the phrase simply played by the user, as if it were played by a professional player of the instrument specified by the user.

[0052] In the embodiment using the inferred tone data 217 output from the sound synthesis LSI 205 in this manner, the inferred tone data 217 can be used to Figure 2 Specifically, for example, the inference tone data 217 output from the sound synthesis LSI 205 can be input to the looper LSI 220 for loop recording, the additionally generated inference tone data 217 can be superimposed on the loop playback sound that is repeatedly played back, and the loop playback sound thus obtained can be used for a performance.

[0053] Figure 4 is a timing diagram of the first embodiment of the loop record / playback process, which is Figure 3 Basic operations of the loop record / playback process executed in the looper LSI 220. In the first embodiment of the loop record / playback process, for ease of understanding, an outline operation will be described for a case where the user executes the loop record / playback process with one phrase and one measure.

[0054] First, when the user steps on Figure 1 When the pedal 105 of FIG10 is pressed, the looper LSI 220 performs the following loop recording / playback processing. In the loop recording / playback processing, a first user operation (performance) is performed in which the user specifies a pitch sequence of a phrase to be repeatedly played back (loop playback) using the keyboard 101 of FIG10. The pitch sequence includes a plurality of pitches with different timings. For example, in Figure 4In the phrase from time t0 (first timing) to time t1 (second timing) in (a), the pitch sequence of the phrase (hereinafter, referred to as "first phrase data") is input from the keyboard 101 via the key scanner 206 and the CPU 201 to the sound synthesis LSI 205 as pitch data 215 (hereinafter, this input is referred to as "first input").

[0055] When the sound synthesis LSI 205 receives data indicating a musical instrument and the first input (pitch data 215 including a pitch sequence of a phrase) from the CPU 201 in advance, for example, from time t1 to time t2 in FIG. (b), the sound synthesis LSI 205 synthesizes inference tone data 217 of the phrase accordingly and outputs it to the looper LSI 220. The sound synthesis LSI 205 generates inference tone data 217 based on the inference tone data to be generated later in FIG. Figure 6 The acoustic feature data 617 output by the trained acoustic model unit 606 described in the above outputs the inferred musical tone data 217 of the musical phrase (hereinafter referred to as "first phrase inferred musical tone data 217"). The inferred musical tone data 217 includes a performance expression sound representing a sound corresponding to a performance technique that has not been performed by the user. The performance technique that has not been performed by the user refers to, for example, a sound performance technique of a performer reproducing a musical phrase performance (hereinafter referred to as "first phrase performance") (including legato) on a musical instrument.

[0056] exist Figure 3 Circulator LSI 220, from Figure 4 From time t1 to time t2 in (b), the loop recording unit 303 sequentially records (stores) the data from the sound synthesis LSI 205 based on the data from the first loop storage area 301 (Area1). Figure 4 The first phrase inference tone data 217 of the phrase is outputted as a result of the first input of the phrase from time t0 to time t1 in (a).

[0057] exist Figure 3 In the looper LSI 220, for example, the loop playback unit 304 outputs (hereinafter, this output is referred to as "first output") the first phrase inference tone data 217 (first sound data, or Figure 4 The first phrase inference tone data 217 is recorded in the first loop storage area 301 (Area1) as the loop playback sound 310 from Figure 4 (c) is repeated from time t1 to time t2, from time t2 to time t3, from time t4 to time 5, and so on. Figure 2 The first output of the circulator LSI 220 repeatedly outputs as loop playback inference music data (first sound data) 222 is output from a speaker (not shown) via the mixer 213, the D / A converter 211, and the amplifier 214.

[0058] Then, in Figure 4 When the playback of the first output is repeated as shown in (c), the user steps on the Figure 1 On the pedal 105, for example, at the start timing of the phrase (for example, Figure 4 After stepping on the pedal 105, the user performs a second user operation (performance) to designate a pitch sequence of another phrase to use. Figure 1 The keyboard 101 is played back in a loop. The pitch sequence includes multiple pitches with different timings. For example, Figure 4 In (d) of the phrase from time t4 (third timing) to time t5 (fourth timing), the pitch sequence of the phrase (hereinafter referred to as "second phrase data") is input from the keyboard 101 via the key scanner 206 and the CPU 201 to the sound synthesis LSI 205 as pitch data 215 (hereinafter, this input is referred to as "second input").

[0059] When the sound synthesis LSI 205 receives data indicating a musical instrument and a second input (pitch data 215 including a pitch sequence of a phrase) from the CPU 201 in advance, for example, Figure 4 From time t5 to time t6 in (e), the sound synthesis LSI 205 synthesizes the inference tone data (second sound data) 217 ​​of the phrase accordingly and outputs it to the looper LSI 220. As with the first phrase inference tone data 217, the sound synthesis LSI 205 synthesizes the inference tone data (second sound data) 217 ​​from later Figure 6 The inferred musical tone data 217 of the musical phrase (hereinafter referred to as "second musical phrase inferred musical tone data 217") is output based on the acoustic feature data 617 output by the trained acoustic model unit 606 described in

[15] . The inferred musical tone data 217 includes a performance expression sound corresponding to a performance technique not performed by the user. The performance technique not performed by the user refers to, for example, a reproduction of another musical phrase performance by a performer on a musical instrument (hereinafter referred to as "second musical phrase performance"), including a legato performance technique.

[0060] In other words, based on the acoustic feature data 617 output from the trained acoustic model unit 606, the processor 201 of the electronic musical instrument 100 is configured to output musical tones 217 to which effects corresponding to various performance techniques are applied, even without detecting user playing operations corresponding to performance techniques.

[0061] exist Figure 3 Circulator LSI 220, from Figure 4 From time t5 to time t6 in (e), the mixer 307 will Figure 4The second phrase inferred tone data 217 of the phrase output from the sound synthesis LSI 205 is the same as the second phrase input from time t4 to time t5 in (d). Figure 4 The first output of the first phrase inference tone data from time t4 to time t5 in (c) is mixed, and the first output is the loop playback sound 310 input to the mixer 307 as the loop playback sound delay output 311 from the loop playback unit 304 via the phrase delay unit 305. Figure 4 From time t5 to time t6 in (e), the loop recording unit 303 sequentially records (stores) the superimposed first output and second output as described above, for example, in the second loop storage area 302 (Area2).

[0062] exist Figure 3 In the looper LSI 220, the loop playback unit 304 outputs (hereinafter, this output is referred to as "second output") the second phrase inference tone data (second sound data) 217 ​​superimposed on the first phrase inference tone data (first sound data) 217, and the second phrase inference tone data is recorded in the second loop storage area 302 (Area2) as the loop playback sound 310, for example, from Figure 4 (f) from time t5 to time t6, from time t6 to time t7, from time t7 to time t8, and so on. Figure 2 The repeated sound sequence in which the first output and the second output are superimposed, output from the circulator LSI 220 , is output from a speaker (not shown) via the mixer 213 , the D / A converter 211 , and the amplifier 214 as loop-playback inference tone data 222 .

[0063] If loop recording of a phrase is to be further superimposed, similar processing may be performed on a new first input (which was once the second input) and a new second input.

[0064] In this manner, according to the first embodiment of the loop recording / reproduction processing performed by the looper LSI 220, only by the user inputting a phrase to designate a pitch sequence as first phrase data and additionally as second phrase data with the second phrase data superimposed on the first phrase data, the first phrase data and the second phrase data can be converted into each inference tone data 217 to reproduce the musical instrument's phrase performance by the player using the sound synthesis LSI 205. Thus, a looped phrase sound sequence including performance expression sounds representing sounds corresponding to performance techniques not performed by the user (such as articulations including legato) can be output as the loop-played back inference tone data 222.

[0065] In the embodiment of the loop recording / playback processing, when the first phrase inference tone data (first sound data) 217 ​​and the second phrase inference tone data (second sound data) 217 ​​are superimposed, the beat of each superimposed phrase may not be synchronized. Figure 3 In the looper LSI 220 of the present embodiment shown, the beat extraction unit 306 is configured to extract the beat timing from the loop playback sound 310 output by the loop playback unit 304 as beat data 221, to output the beat data 221 to the sound synthesis LSI 205, thereby performing quantization processing.

[0066] Figure 5 The quantization process is shown in the figure. Figure 5 In (a-1), the four filling blocks in the phrase from time t0 to time t1 represent the first phrase data ( Figure 4 The “first input” in (a) is the pitch sequence of four notes, and the user’s keystroke timing for each note in the phrase from time t0 to time t1 is expressed as Figure 4 The corresponding keys in the keyboard 101 in (a) are schematically shown. Figure 4 On the other hand, if Figure 5 1 shows first phrase inference tone data (first sound data) output from the sound synthesis LSI 205 by inputting the first phrase data into the sound synthesis LSI 205 as the pitch data 215. A1, a3, a5, and a7 correspond to notes played in the beat, while a2, a4, and a6 schematically show notes played in a legato manner, which are not in the original user performance.

[0067] Taking into account that the performer's performance technique is reproduced, the note-on timings of a1, a3, a5, and a7 do not always coincide with each beat.

[0068] exist Figure 5 In (a-2), the four filling blocks in one phrase from time t4 to time t5 represent the second phrase data ( Figure 4 The "second input" in (d) is a four-note pitch sequence, and the user's keystroke timing for each note in the phrase from time t4 to time t5 is used in Figure 4 The corresponding keys in the keyboard 101 in (d) are schematically shown. Figure 4 When the second phrase inference tone data 217 generated by inputting the second phrase data into the sound synthesis LSI 205 is superimposed on Figure 4 When the first phrase inference tone data 217 is played in the loop recording / playback process described in , since the two pitch sequences are generally different from each other, the beats of the two inference tone data 217 may not be synchronized.

[0069] Therefore, according to Figure 3 In the looper LSI 220 of the present embodiment shown, the beat extraction unit 306 is configured to extract the beat timing from the loop playback sound 310, which is played by Figure 5 The first phrase inference tone data 217 in (b-1) is generated, and the beat data 221 Figure 5 The note-on timings of a1, a3, a5, and a7 in (b-1) of FIG. 3 can be extracted by detecting four power peaks of the waveform of the looped playback sound 310, for example. The beat data 221 represents the beats in the phrase extracted from the looped playback sound 310 and is input to the MIDI data to be described later. Figure 6 The oscillation generating unit 609 in the sound model unit 608 in the sound synthesis LSI 205. Figure 6 , the oscillation generating unit 609 performs so-called quantization processing, in which, for example, each pulse is adjusted according to the beat data 221 when generating a pulse train that is periodically repeated at the basic frequency (F0) included in the sound source data 619.

[0070] Under such control, the Figure 6 The sound model unit 608 in the sound synthesis LSI 205 is Figure 4 The second phrase inference tone data 217 is generated during the phrase from time t5 to time t6 in (e) so that the note-on timings of the notes in the beat of the second phrase inference tone data 217 can be synchronized with the note-on timings of the notes a1, a3, a5, and a7 of the first phrase inference tone data 217, as shown in FIG. Figure 5 Therefore, from the later described Figure 6 The inferred tone data 217 output by the sound model unit 608 can be synchronized with the loop playback sound 310 already generated in the looper LSI 220, so that the loop playback sound 310 will be less discordant even when overdubbing.

[0071] That is, the processor 205 generates the second sound data by synchronizing the note-on timing of the first note in the first sound data with the note-on timing of the second note in the second phrase data, wherein the note-on timing of the first note is not always consistent with the beat, or generates the second sound data by matching the duration of the first note in the first sound data with the duration of the second note in the second phrase data.

[0072] Figure 6 6 is a block diagram showing a configuration example of the voice synthesis unit 602 and the voice training unit 601 in this embodiment. Figure 2The functions performed by the sound synthesis LSI 205 are built into the electronic keyboard instrument 100.

[0073] The sound synthesis unit 602 recognizes each phrase (measure) based on a rhythm setting described later. Figure 1 The keys of the keyboard 101 are touched via Figure 2 The key scanner 206 receives pitch data 215 including a pitch sequence instructed from the CPU 201 to synthesize and output inferred musical tone data 217. In response to the operation of a key (operator) on the keyboard 101, the processor of the sound synthesis unit 602 inputs the pitch data 215 including the pitch sequence of the musical phrase associated with the key into the trained acoustic model of the musical instrument selected by the user in the trained acoustic model unit 606. The processor of the sound synthesis unit 602 performs a process of outputting the inferred musical tone data 217, which reproduces the instrument performance sound of the musical phrase based on the spectrum data 618 and sound source data 619 output by the trained acoustic model unit 606 according to the input.

[0074] For example, Figure 6 As shown, the sound training unit 601 can be implemented as a Figure 1 Alternatively, although not in Figure 6 As shown, but if Figure 2 If the sound synthesis LSI 205 has spare processing capabilities, the sound training unit 601 can be built into the electronic keyboard instrument 100 as a function executed by the sound synthesis LSI 205.

[0075] For example, based on the following non-patent document 1 "Statistical Parametric Speech Synthesis Based on Deep Learning" Figure 2 The voice training unit 601 and the voice synthesis unit 602.

[0076] Non-patent document 1: Kei Hashimoto and Shinji Takaki, “Statistical Parametric Speech Synthesis Based on Deep Learning,” Journal of the Acoustical Society of Japan, Vol. 73, No. 1 (2017), pp. 55-62.

[0077] like Figure 6As shown, for example, the function performed by the external server 600 Figure 2 The sound training unit 601 includes a training acoustic feature extraction unit 604 and a model training unit 605.

[0078] In the sound training unit 601, for example, a sound recording obtained by playing a music piece of a certain genre on a musical instrument is used as a training performance data set 612, and text data of a pitch sequence of a phrase of the music piece is used as a training score data set 611, which is a training pitch data set.

[0079] The training acoustic feature extraction unit 604 is configured to load and analyze a training performance dataset 612 recorded using a microphone or the like, for example, by a professional performer playing the pitch sequences of the musical phrases included in the training musical score dataset 611 on a musical instrument, and load text data of the pitch sequences of the musical phrases included in the training musical score dataset 611. The training acoustic feature extraction unit 604 then extracts and outputs a training acoustic feature sequence 614 representing the sound features in the training performance dataset 612.

[0080] The model training unit 605 is configured to estimate the acoustic model using machine learning so that the conditional probability of the training acoustic feature sequence 614 given the training music score dataset 611 and the acoustic model is maximized according to equation (1) described in the following non-patent document 1. In other words, the relationship between the instrument sound feature sequence as text data and the acoustic feature sequence as sound is expressed using a statistical model called an acoustic model.

[0081] [Mathematical expression 1]

[0082]

[0083] In equation (1), arg max refers to an operation that returns the parameter written below it so as to maximize the value of the function written on the right.

[0084] The following symbols represent the training music score dataset 611.

[0085] [Mathematical expression 2]

[0086] l

[0087] The following symbols represent the acoustic model.

[0088] [Mathematical expression 3]

[0089] λ

[0090] The following symbols represent the training acoustic feature sequence 614.

[0091] [Mathematical expression 4]

[0092] o

[0093] The following symbols represent the probability that a training acoustic feature sequence 614 will be generated.

[0094] [Mathematical expression 5]

[0095] P(o|l,λ)

[0096] The following notation represents the acoustic model such that the probability that the training acoustic feature sequence 614 will be generated is maximized.

[0097] [Mathematical expression 6]

[0098]

[0099] The model training unit 605 is configured to output model parameters as training results 615 , the model parameters representing the acoustic model calculated according to equation (1) using machine learning.

[0100] For example, Figure 6 As shown, in Figure 1 Before the electronic keyboard instrument 100 is shipped, the learning result 615 (model parameter) is stored in the electronic keyboard instrument 100. Figure 2 When the electronic keyboard instrument 100 is turned on, the training result 615 can be read from the ROM 202 of the control system shown in FIG. Figure 2 The ROM 202 is loaded into the trained acoustic model unit 606 described later in the sound synthesis LSI 205. Or, for example, as Figure 6 As shown, in response to a user operation on the second switch panel 103 of the electronic keyboard instrument 100, the training result 615 can be downloaded from the Internet (not shown) or a network using a universal serial bus (USB) cable, etc. through the network interface 219 to the trained acoustic model unit 606 described later in the sound synthesis LSI 205.

[0101] The sound synthesis unit 602 is a function performed by the sound synthesis LSI 205, and includes a trained acoustic model unit 606 and an acoustic model unit 608. The sound synthesis unit 602 is configured to perform statistical sound synthesis processing, in which the musical tone data 217 is inferred to correspond to the pitch data 215 of the text data including the pitch sequence of the musical phrase by making predictions using a statistical model called an acoustic model in the trained acoustic model unit 606.

[0102] The trained acoustic model unit 606 receives the pitch data 215 of the musical phrase and outputs a corresponding predicted acoustic feature sequence 617. In other words, the trained acoustic model unit 606 is configured to estimate the estimated value of the acoustic feature sequence 617 so that the conditional probability of the acoustic feature sequence 617 as acoustic feature data, given the pitch data 215 input from the keyboard 101 via the key scanner 206 and the CPU 201 and the acoustic model set as the learning result 615 using machine learning in the model training unit 605, is maximized according to Equation (2) described in the following non-patent document 1.

[0103] [Mathematical expression 7]

[0104]

[0105] In equation (2), the following symbols represent the pitch data 215 input from the keyboard 101 via the key scanner 206 and the CPU 201.

[0106] [Mathematical expression 8]

[0107] l(2-1)

[0108] The following symbols represent the acoustic model set as the learning result 615 using machine learning in the model training unit 605.

[0109] [Mathematical expression 9]

[0110]

[0111] The following symbols represent the acoustic feature sequence 617 as acoustic feature data.

[0112] [Mathematical expression 10]

[0113] o(2-3)

[0114] The following symbols represent the probability that the acoustic feature sequence 617 will be generated as acoustic feature data.

[0115] [Mathematical expression 11]

[0116]

[0117] The following symbols represent estimated values ​​of the acoustic feature sequence 617 such that the probability that the acoustic feature sequence 617 will be generated as acoustic feature data is maximized.

[0118] [Mathematical expression 12]

[0119]

[0120] The sound model unit 608 receives the acoustic feature sequence 617 and generates the inferred tone data 217 corresponding to the pitch data 215 including the pitch sequence specified by the CPU 201. The inferred tone data 217 is input to Figure 2 circulator 220.

[0121] The acoustic features represented by the trained acoustic feature sequence 614 and the acoustic feature sequence 617 include spectral data and sound source data. The spectral data models the mechanism of sound generation or resonance of the musical instrument, and the sound source data models the oscillation mechanism of the musical instrument. As spectral data (spectral parameters), Mel-frequency cepstrum, line spectral pair (LSP), etc. can be used. As sound source data, the fundamental frequency (F0) and power representing the pitch frequency of the musical instrument sound can be used. The sound model unit 608 includes an oscillation generation unit 609 and a synthesis filter unit 610. The oscillation generation unit 609 models the oscillation mechanism of the musical instrument. The oscillation generation unit 609 sequentially receives the sequence of sound source data 619 output from the trained acoustic model unit 606 and generates a sound signal, for example, composed of a pulse sequence that periodically repeats at the fundamental frequency (F0) and power (in the case of voiced notes) included in the sound source data 619, white noise having the power included in the sound source data 619 (in the case of unvoiced notes), or a mixture thereof. The synthesis filter unit 610 simulates the mechanism of sound generation or resonance of a musical instrument. The synthesis filter unit 610 sequentially receives the sequence of spectrum data 618 output from the trained acoustic model unit 606 and forms a digital filter to model the mechanism of sound generation or resonance of the musical instrument. The synthesis filter unit 610 then uses the sound signal input from the oscillation generating unit 609 as an oscillator signal to generate and output the inferred musical tone data 217 as a digital signal.

[0122] The sampling rate of the training performance dataset 612 is, for example, 16 kHz. If the mel-frequency parameters obtained by the mel-frequency cepstrum analysis are used as the spectral parameters included in the training acoustic feature sequences 614 and 617, the 1st to 24th MFCCs are obtained, for example, with a frame shift of 5 milliseconds, a frame size of 25 milliseconds, and a Blackman window as the window function.

[0123] The inference musical tone data 321 output from the sound synthesis unit 602 is input to Figure 2 circulator 220.

[0124] Next, we will explain Figure 6A first embodiment of the statistical sound synthesis processing performed by the sound training unit 601 and the sound synthesis unit 602 is provided. In the first embodiment of the statistical sound synthesis processing, a hidden Markov model (HMM) described in the non-patent document 1 and the non-patent document 2 cited below is used as the acoustic model represented by the training result 615 (model parameter) set in the trained acoustic model 606.

[0125] Non-patent document 2: Shinji Sako, Keijiro Saino, Yoshihiko Nankaku, Keiichi Tokuda, and Tadashi Kitamura, “A trainable singing voice synthesis system capable of representing personal characteristics and singing styles,” Information Processing Society of Japan (IPSJ) Technical Report, Music and Computer (MUS), Vol. 2008, No. 12 (2008), pp. 39-44.

[0126] In the first embodiment of the statistical sound synthesis process, when a musical instrument sound is given in a pitch sequence of a musical phrase by a user's performance, an HMM acoustic model is trained based on how the characteristic parameters of the musical instrument sound, such as the sound source and sound generation or resonance characteristics of the musical instrument, change over time. More specifically, the HMM acoustic model is trained based on the note sound (the sound of the note played in the musical score) from the training instrument data ( Figure 6 The spectrum, fundamental frequency (pitch) and its time structure are modeled by using the training music score dataset 611).

[0127] First, we will describe Figure 6 The sound training unit 601 uses the HMM model for processing. By inputting the training score data set 611 and the training acoustic feature sequence 614 output from the training acoustic feature extraction unit 604, the model training unit 605 in the sound training unit 601 trains the HMM acoustic model based on equation (1) with maximum likelihood. As described in Non-Patent Document 1, the likelihood function of the HMM acoustic model is expressed by the following equation (3).

[0128] [Mathematical expression 13]

[0129]

[0130] In equation (3), the following symbols represent acoustic features in frame t.

[0131] [Mathematical expression 14]

[0132] o t

[0133] T represents the frame number.

[0134] The following symbols represent the state sequence of the HMM acoustic model.

[0135] [Mathematical expression 15]

[0136] q=(q1,…,q T )

[0137] The following symbols represent the state number of the HMM acoustic model in frame t.

[0138] [Mathematical expression 16]

[0139] q t

[0140] The following symbol represents the transition from state q t-1 To state q t The state transition probability.

[0141] [Mathematical expression 17]

[0142]

[0143] The following symbols are with mean vector μ qt and the covariance matrix Σ qt The normal distribution of , and represents the state q t The output probability distribution of . Note that μ qt and Σ qt The t in the string is the subscript of q.

[0144] [Mathematical expression 18]

[0145]

[0146] The HMM acoustic model is trained efficiently based on the maximum likelihood criterion using the Expectation Maximization (EM) algorithm.

[0147] The spectral data of musical instrument sounds can be modeled using a continuous HMM. However, since the logarithmic fundamental frequency (F0) is a variable-dimensional time series signal, it takes continuous values ​​in the voiced segments (segments with pitch) of the musical instrument sound, but does not take values ​​in the unvoiced segments (segments without pitch, such as breathing sounds), and cannot be directly modeled using an ordinary continuous or discrete HMM. Therefore, a multi-space probability distribution HMM (MSD-HMM) is used, which is an HMM based on a multi-space probability distribution compatible with variable dimensions, to model the Mel-frequency cepstrum (spectral parameter) as a multivariate Gaussian distribution, and to model the voiced sound of the musical instrument sound with the logarithmic fundamental frequency (F0) as a Gaussian distribution in a one-dimensional space, and at the same time, to model the unvoiced sound of the musical instrument sound as a Gaussian distribution in a zero-dimensional space.

[0148] The acoustic features (pitch, duration, start / end timing, beat emphasis, etc.) of the note sounds that make up the sound of an instrument are known to vary due to the influence of various factors, even if the notes (e.g., the pitch of the notes) are the same. These factors that affect the acoustic features of the note sounds are called context. In the statistical sound synthesis processing of the first embodiment, in order to accurately model the acoustic features of the note sounds of the instrument, an HMM acoustic model (context-dependent model) that takes context into account can be used. Specifically, the training score dataset 611 can consider not only the pitch of the note sounds, but also the pitch sequence of the notes in the phrases, instruments, etc. In the model training unit 605, context clustering based on decision trees can be used to effectively process the combination of contexts. In this clustering, a set of HMM acoustic models is divided into a tree structure using a binary tree, so that HMM acoustic models with similar contexts are grouped into clusters. Each node in the tree has a question for dividing the context into two groups, such as "Is the pitch of the previous note sound x?", "Is the pitch of the next note sound y?", and "Is the instrument z?" Each leaf node has a training result 615 (model parameters) corresponding to a specific HMM acoustic model. For any combination of contexts, by traversing the tree according to the question at the node, one of the leaf nodes can be reached, and the training result 615 (model parameters) corresponding to the leaf node can be selected. By selecting an appropriate decision tree structure, an HMM acoustic model (context-dependent model) with high accuracy and high generalization ability can be estimated.

[0149] Figure 7 This is a diagram for explaining the HMM decision tree in the first embodiment of the statistical sound synthesis process. For each note sound depending on the context, the state of the note sound is, for example, the same as that determined by Figure 7The HMM consisting of three states 701, #1, #2, and #3, shown in (a) of FIG. The incoming or outgoing arrows of each state represent state transitions. For example, state 701 (#1) models the note-on state of a note sound, state 701 (#2) models the middle state of a note sound, and state 701 (#3) models the note-off state of a note sound.

[0150] Depending on the duration of the note sound, Figure 7 The duration of each state 701 (#1) to (#3) shown in the HMM of (a) is calculated using Figure 7 The state duration model in (b) is determined. Figure 6 The model training unit 605 is obtained by Figure 6 The state duration decision tree 702 is generated by learning from the training music score dataset 611 for determining the state duration, where the training music score dataset 611 corresponds to the context of a large number of note sound sequences based on musical phrases, and the state duration decision tree 702 is set as the training result 615 in the trained acoustic model unit 606 in the sound synthesis unit 602.

[0151] in addition, Figure 6 The model training unit 605 generates a Mel-cepstrum parameter decision tree 703 to determine Mel-cepstrum parameters by learning from a training acoustic feature sequence 614, wherein the training acoustic feature sequence 614 corresponds to a large number of note sound sequences based on a phrase with respect to Mel-frequency cepstrum parameters, and the training acoustic feature sequence 614 is, for example, obtained by Figure 6 The training acoustic feature extraction unit 604 is obtained from Figure 6 The mel-frequency cepstrum parameter decision tree 703 is extracted from the training performance data set 612. Then, the model training unit 605 sets the generated mel-frequency cepstrum parameter decision tree 703 as the training result 615 in the trained acoustic model unit 606 in the sound synthesis unit 602.

[0152] also, Figure 6 The model training unit 605 generates a logarithmic fundamental frequency decision tree 704 for determining the logarithmic fundamental frequency (F0) by learning from the training acoustic feature sequence 614, wherein the training acoustic feature sequence 614 corresponds to a large number of note sound sequences based on the musical phrases about the logarithmic fundamental frequency (F0) and is, for example, determined by Figure 6 The training acoustic feature extraction unit 604 is obtained from Figure 6The model training unit 605 then sets the generated logarithmic fundamental frequency decision tree 704 as a training result 615 in the trained acoustic model unit 606 in the sound synthesis unit 602. Note that, as described above, using the MSD-HMM compatible with variable dimensionality, voiced segments with a logarithmic fundamental frequency (F0) are modeled as a Gaussian distribution in a one-dimensional space, while unvoiced segments are modeled as a Gaussian distribution in a zero-dimensional space. In this way, the logarithmic fundamental frequency decision tree 704 is generated.

[0153] Although Figure 7 Not shown, but Figure 6 The model training unit 605 can be configured to generate a decision tree for determining context regarding the accents of note sounds (e.g., beat accents) and the like by learning from a training music score dataset 611. The training music score dataset 611 corresponds to the context of a large number of note sound sequences on a phrase basis. The model training unit 605 can then set the generated decision tree as a training result 615 in the trained acoustic model unit 606 in the sound synthesis unit 602.

[0154] Next, the following will be described. Figure 6 The trained acoustic model 606 loads the pitch data 215 in the context of the pitch sequence in the phrase of the instrument sound of the musical instrument, which is input from the keyboard 101 via the key scanner 206 and the CPU 201, to generate the pitch data 215 by referring to each context. Figure 7 The HMMs are connected by using the decision trees 702, 703, and 704, etc. Then, the trained acoustic model 606 predicts the acoustic feature sequence 617 (spectral data 618 and sound source data 619) so that the output probability is maximized using each connected HMM.

[0155] At this time, the trained acoustic model unit 606 estimates the estimated value of the acoustic feature sequence 617 (symbol (2-5)) so that the conditional probability (symbol (2-4)) of the acoustic feature sequence 617 (symbol (2-3)) is maximized given the pitch data 215 (symbol (2-1)) input from the keyboard 101 through the key scanner 206 and the CPU 201 and the acoustic model (symbol (2-2)) of the training result 615 set according to equation (2) using the machine learning in the model training unit 605. Figure 7 The state sequence (4-1) estimated by the state duration model in (b), equation (2) is approximated as equation (4) described in the following non-patent document 1.

[0156] [Mathematical expression 19]

[0157]

[0158] [Mathematical expression 20]

[0159]

[0160] [Mathematical expression 21]

[0161]

[0162] The left sides of equations (4-2) and (4-3) above are the mean vector and covariance matrix of the state (4-4) below, respectively.

[0163] [Mathematical expression 22]

[0164]

[0165] Using the instrument sound feature sequence l, the mean vector and covariance matrix are calculated by traversing each decision tree set in the trained acoustic model 606. According to equation (4), the estimated value of the acoustic feature sequence 617 (symbol (2-5)) is obtained using the mean vector of the above equation (4-2), which is a discontinuous sequence that changes in a step-like manner when the state transitions. If the synthesis filter unit 610 synthesizes the inference musical sound data 217 from this discontinuous acoustic feature sequence 617, low-quality or unnatural instrument sounds are produced. Therefore, in the first embodiment of the statistical sound synthesis process, an algorithm for generating a training result 615 (model parameter) that takes into account dynamic features can be adopted in the model training unit 605. If the acoustic feature sequence in frame t (equation (5-1) below) consists of static features and dynamic features, then the change of the acoustic feature sequence (equation (5-2) below) over time is represented by equation (5-3) below.

[0166] [Mathematical expression 23]

[0167]

[0168] o=Wc (5-3)

[0169] In the above equation (5-3), W is a matrix used to obtain the acoustic feature sequence o including dynamic features from the static feature sequence of the following equation (6-4).

[0170] [Mathematical expression 24]

[0171]

[0172] The model training unit 605 solves the above equation (4), as expressed by the following equation (6) with the above equation (5-3) as constraints.

[0173] [Mathematical expression 25]

[0174]

[0175] The left side of equation (6) above is a static feature sequence that maximizes the output probability under the constraints of the dynamic features. By considering the dynamic features, the discontinuity at the state boundary can be resolved to obtain a smoothly changing acoustic feature sequence 617. Therefore, high-quality inference musical tone data 217 can be generated in the synthesis filter unit 610.

[0176] Next, explain Figure 6 The present invention is a second embodiment of the statistical sound synthesis processing performed by the sound training unit 601 and the sound synthesis unit 602. In the second embodiment of the statistical sound synthesis processing, in order to predict the acoustic feature sequence 617 from the pitch data 215, a deep neural network (DNN) is used to implement the trained acoustic model unit 606. Accordingly, the model training unit 605 in the sound training unit 601 is configured to learn the model parameters of the nonlinear transformation function representing the neurons in the DNN from the instrument sound features (training music score data set 611) to the acoustic features (training acoustic feature sequence 614) and output the model parameters to the DNN of the trained acoustic model unit 606 in the sound synthesis unit 602 as the learning result 615.

[0177] Typically, for example, acoustic features are calculated for each frame with a width of 5.1 milliseconds, and instrument sound features are calculated for each note. Therefore, the time units of acoustic features and instrument sound features are different. In the first embodiment of the statistical sound synthesis process using the HMM acoustic model, the correspondence between the acoustic features and the instrument sound features is represented by the state sequence of the HMM, and the model training unit 605 is based on Figure 6 The training score data set 611 and the training performance data set 612 are used to automatically learn the correspondence between acoustic features and instrument sound features. In contrast, in the second embodiment of the statistical sound synthesis processing using DNN, since the DNN set in the trained acoustic model unit 606 is a model that represents a one-to-one correspondence between the pitch data 215 as input and the acoustic feature sequence 617 as output, it is impossible to use a pair of input and output data with different time units to train the DNN. For this reason, in the second embodiment of the statistical sound synthesis processing, the correspondence between the acoustic feature sequence in frames and the instrument sound feature sequence in notes is pre-set, and a pair of acoustic feature sequences and instrument sound feature sequences in frames are generated.

[0178] Figure 8The operation of the sound synthesis LSI 205 showing the above correspondence is shown. For example, if a musical instrument note sequence is given, it is in a phrase of a piece of music ( Figure 8 (a)) of the pitch sequence (string) "C3", "E3", "G3", "G3", "G3", "G3" and other musical instrument sound feature sequences, then the musical instrument sound feature sequences are in a one-to-many correspondence relationship ( Figure 8 (a) and (b)) and the acoustic feature sequence with the unit of frame ( Figure 8 (b)). Note that the instrument sound features must be represented as numerical data because they are used as input to the DNN in the trained acoustic model unit 606. For this purpose, as the instrument sound feature sequence, numerical data obtained by concatenating binary (0 or 1) or continuous-valued data for context-related questions (e.g., "Was the previous note x?" and "What is the instrument of the current note y?") is used.

[0179] In the second embodiment of the statistical sound synthesis process, as shown in FIG. Figure 8 As shown by the dotted arrow 801. Figure 6 The model training unit 605 in the sound training unit 601 trains the DNN by sequentially passing the pairs of frames of the training score dataset 611, which is a sequence of notes (pitch sequence) of a musical phrase and corresponds to Figure 8 (a), and the training acoustic feature sequence 614 of the phrase corresponds to Figure 8 (b) becomes the DNN in the trained acoustic model unit 606. Note that the DNN in the trained acoustic model unit 606 includes neurons such as Figure 8 As shown by the gray circle in , it consists of an input layer, one or more hidden layers, and an output layer.

[0180] On the other hand, when synthesizing sound, the note sequence (pitch sequence) of the phrase is used as a frame unit and corresponds to Figure 8 The pitch data 215 of (a) is input to the DNN in the trained acoustic model unit 606. Therefore, the DNN in the trained acoustic model unit 606 outputs an acoustic feature sequence 617 of the phrase in frames, such as Figure 8 As shown by the thick solid arrow 802. Therefore, also in the sound model unit 608, the sound source data 619 and the spectrum data 618 included in the acoustic feature sequence 617 of the phrase and whose unit is a frame are given to the oscillation generating unit 609 and the synthesis filter unit 610 accordingly, thereby performing sound synthesis.

[0181] Therefore, the sound model unit 608 generates the sound model by corresponding to, for example, Figure 8The 225-sample frame indicated by the thick solid arrow 803 outputs the inferred musical tone data 217 of the phrase. Since the frame width is 5.1 milliseconds, one sample corresponds to 5.1 milliseconds / 225≈0.0227 milliseconds. The sampling rate of the inferred musical tone data 217 is therefore 1 / 0.0227≈44 kHz.

[0182] According to the ordinary least squares criterion, Equation (7) below, a DNN is trained using a pair of acoustic features and instrument features (pitch sequence and instrument) of a musical phrase in frame units.

[0183] [Mathematical expression 26]

[0184]

[0185] The following symbols represent the acoustic features in frame t, numbered t.

[0186] [Mathematical expression 27]

[0187] o t

[0188] The following symbols represent the instrument sound features (pitch and instrument) in frame t, numbered t.

[0189] [Mathematical expression 28]

[0190] l t

[0191] The following symbols represent the model parameters of the trained DNN in the acoustic model unit 606.

[0192] [Mathematical expression 29]

[0193]

[0194] The following symbols represent the nonlinear transformation function represented by the DNN. The model parameters of the DNN can be efficiently estimated using backpropagation.

[0195] [Mathematical expression 30]

[0196] g λ (·)

[0197] Taking into account the corresponding relationship with the processing of the model training unit 605 in the statistical sound synthesis represented by the above equation (1), the training of the DNN can be expressed as the following equation (8).

[0198] [Mathematical expression 31]

[0199]

[0200] In the above equation (8), the following equation (9) holds.

[0201] [Mathematical expression 32]

[0202]

[0203] Just like in equations (8) and (9) above, the relationship between acoustic features and instrument sound features (pitch and instrument) can be represented using a normal distribution, as shown in equation (9-1), and the output of the DNN is the mean vector.

[0204] [Mathematics 33]

[0205]

[0206] Generally, in the second embodiment of the statistical sound synthesis process using DNN, a sound feature sequence independent of the instrument sound feature sequence is used. t The covariance matrix of , that is, the covariance matrix common to all frames (Equation (9-2) below).

[0207] [Mathematical expression 34]

[0208]

[0209] If the covariance matrix of the above equation (9-2) is the identity matrix, the above equation (8) represents a training process equivalent to the above equation (7).

[0210] like Figure 8 As shown, the trained DNN in the acoustic model unit 606 is configured to independently predict an acoustic feature sequence 617 for each frame. Therefore, the obtained acoustic feature sequence 617 may include discontinuities that reduce the quality of the synthesized sound. Therefore, in this embodiment, a parameter generation algorithm using dynamic features similar to the first embodiment of the statistical sound synthesis process can be used to improve the quality of the synthesized sound.

[0211] The following uses Figure 2 The inference tone data 217 outputted from the sound synthesis LSI 205 is described in detail for realizing the Figure 2 The loop record / playback process of the second embodiment performed by the looper LSI 220 Figure 1 and Figure 2 More specific operations of the electronic keyboard instrument 100 are shown. In the second embodiment of the loop record / playback process, a plurality of consecutive phrases can be set as a loop section.

[0212] In the second embodiment of the loop recording / playback process, in a state where the loop recording / playback process is not executed (hereinafter referred to as "Mode 0"), once the user steps on the Figure 1, the circulator LSI 220 shifts to the operation of Mode 1 (Mode 1) described below. In Mode 1, when the user Figure 1 When performing a performance on the keyboard 101, specifying the desired pitch sequence for each phrase, Figure 2 The sound synthesis LSI 205 outputs the inferred musical tone data 217, which outputs the added performance expression sound in phrase units with a delay of one phrase (one measure in this embodiment), and the performance expression sound includes sounds corresponding to the performance techniques that the user has not performed, as shown in FIG. Figures 6 to 8 As described. Figure 2 The circulator LSI 220 Figure 3 The loop recording unit 303 shown executes a process of sequentially storing the inferred musical tone data 217 in the first loop storage area 301 via the mixer 307. The inferred musical tone data 217 is output from the sound synthesis LSI 205 for each phrase as described above, based on the pitch sequence performed by the user on a phrase basis. While the user calculates a plurality of phrases (measures), the user performs the above-mentioned performance based on the rhythmic sound emitted from the speaker (not shown) via, for example, the mixer 213, the D / A converter 211, and the amplifier 214 of the sound module LSI 204, thereby causing the circulator LSI 220 to execute the operation of Mode 1.

[0213] Thereafter, when the user steps on the pedal 105 again, the circulator LSI 220 switches to the operation of Mode 2 (Mode 2) which will be described below. In Mode 2, Figure 3 The loop playback unit 304 is configured to sequentially load the inference musical tone data 217 of a plurality of phrase (measure) segments (hereinafter referred to as "loop segments") stored in the first loop storage area 301 as loop playback sounds 310. The loop playback sounds 310 are input to Figure 2 The mixer 213, as Figure 2 The loop playback reasoning tone data 222 is played back and emitted from a speaker (not shown) through the digital-to-analog converter 211 and the amplifier 214. The user further Figure 1 The following performance is performed on the keyboard 101 of the loop section, and for each phrase in the loop section, the desired pitch sequence corresponding to the musical tone that the user wants to loop record is specified and superimposed on the loop playback sound 310. As a result, Figure 2 The sound synthesis LSI 205 outputs the inference tone data 217 to which rich musical expression is added in phrase units having a delay of one phrase, similarly to the case of Mode 1. Figure 3The mixer 307 is configured to mix the inference tone data 217 input from the sound synthesis LSI 205 in phrase units delayed by one phrase relative to the user's performance with the loop playback sound delay output 311 obtained by delaying the loop playback sound 310 output from the loop playback unit 304 by one phrase in the phrase delay unit 305, and input the mixed data to the loop recording unit 303. The loop recording unit 303 is configured to perform sequential storage of the mixed (so-called overdubbing) inference tone data to the loop recording unit 303. Figure 3 The second loop storage area 302 is processed. When the overdubbing operation reaches the end of the loop section, the loop playback unit 304 switches the loading source of the loop playback sound 310 from the end of the loop in the first loop storage area 301 to the beginning of the loop in the second loop storage area 302. The loop recording unit 303 is configured to switch the recording destination of the inference musical tone data from the end of the second loop storage area 302 to the beginning of the loop in the first loop storage area 301. Furthermore, when the operation reaches the end of the loop section, the loop playback unit 304 again switches the loading source of the loop playback sound 310 from the end of the loop in the second loop storage area 302 to the beginning of the loop in the first loop storage area 301. The loop recording unit 303 is configured to switch the recording destination of the inference musical tone data from the end of the first loop storage area 301 to the beginning of the loop in the second loop storage area 302. Repeating the switching control operation allows the user to generate the loop playback sound 310 of the loop section while sequentially superimposing the inference musical tone data 217 obtained based on the user's performance on the loop playback sound 310 of the loop section.

[0214] Thereafter, when the user steps on the pedal 105 again, the circulator LSI 220 switches to the operation of Mode 3 (Mode 3) to be described below. In Mode 3, Figure 3 The loop playback unit 304 is configured to repeatedly play back the loop playback sound 310 in the loop section starting from the last recording area of ​​the first loop storage area 301 or the second loop storage area 302, and output it as the inference tone data 222. The repeatedly played tone data 217 is output from Figure 2 The mixer 213 is output from the speaker (not shown) via the digital-to-analog converter 211 and the amplifier 214. In this way, even when the user is Figure 1 When a monotonous performance in which the pitch of each note is specified in units of phrases is performed on the keyboard 101 of the user, a loop playback sound 310 having rich musical expression generated by the sound synthesis LSI 205 can be played back (the dynamics of the output sound can fluctuate according to the performance, sounds or acoustic effects can be added, and unplayed notes can be supplemented). At this time, when the user further performs the performance, Figure 2 The sound module LSI 204 outputs the musical sound data based on the performance output. Figure 2The mixed sound is mixed with the loop playback sound 310 in the mixer 213, and the mixed sound can be emitted from the speaker (not shown) through the digital-to-analog converter 211 and the amplifier 214, so that the ensemble of loop playback and user performance can be achieved.

[0215] Thereafter, once when the user steps on the pedal 105 again in the loop playback state of Mode 3, the looper LSI 220 returns to the operation of Mode 2 and can further perform overdubbing.

[0216] If the user holds down pedal 105 in Mode 2 while recording overdubs, looper LSI 220 cancels the last recorded loop recording, switches to Mode 3, and returns to the previous loop recording state. Furthermore, if the user steps on pedal 105 again, looper LSI 220 returns to Mode 2 and continues recording overdubs.

[0217] In the state of Mode 1, Mode 2, or Mode 3, when the user quickly steps on the pedal 105 twice, the looper LSI 220 switches to the stop state of Mode 0 to end the loop recording / playback.

[0218] Figure 9 This is a main flow chart showing an example of control processing of an electronic musical instrument in the second embodiment of loop recording / playback processing. Figure 2 The CPU 201 loads the control processing program from the ROM 202 to the RAM 203 to operate.

[0219] After executing the initialization process (step S901 ), the CPU 201 repeatedly executes a series of processes from step S902 to step S907 .

[0220] In the repetitive processing, the CPU 201 first performs a switching process (step S902). The CPU 201 performs a switching process based on the Figure 2 The interrupt of the key scanner 206 is executed with Figure 1 The processing corresponds to the switch operation on the first switch panel 102 or the second switch panel 103.

[0221] Then the CPU 201 based on the Figure 2 The interrupt of the key scanner 206 is used to determine and process Figure 1 In the keyboard processing, the CPU 201 sends a key press or release operation to the user according to whether any key of the keyboard 101 is operated (step S903). Figure 2The sound module LSI 204 outputs the musical tone generation control data 216 for instructing note-on or note-off. In addition, in the keyboard processing, the CPU 201 performs processing for sequentially storing the pitches of the pressed keys in the phrase buffer, which is an array variable on the RAM 203, so as to output the pitch sequence to the sound synthesis LSI 205 in phrase units in the TickTime interrupt processing described later.

[0222] The CPU 201 then executes display processing of the processed data, and the processed data will be displayed on the Figure 1 on the LCD 104, and through Figure 2 The LCD controller 208 displays data on the LCD 104 (step S904). The data displayed on the LCD 104 is, for example, a musical score corresponding to the played inference tone data 217 and various setting contents.

[0223] Then the CPU 201 executes the circulator control process (step S905). In this process, the CPU 201 executes the circulator control process (step S905). Figure 2 Circulator control processing (described later) of the circulator LSI 220 Figure 14 and Figure 15 Flowchart processing).

[0224] Subsequently, the CPU 201 executes sound source processing (step S906 ). In the sound source processing, the CPU 201 executes control processing such as envelope control of musical tones during sound generation in the sound module LSI 204 .

[0225] Finally, the CPU 201 determines whether the user presses the power off switch (not shown) to turn off the power (step S907). When the determination in step S907 is "No", the CPU 201 returns to the processing of step S902. When the determination in step S907 is "Yes", the CPU 201 ends Figure 9 The control process shown in the flowchart of FIG. 1 is executed, and the power of the electronic keyboard instrument 100 is turned off.

[0226] Figure 10A It shows Figure 9 Flowchart of a detailed example of the initialization process of step S901 in FIG. Figure 10B It is shown in Figure 9 The switching process of step S902 in the Figure 11 Flowchart of a detailed example of the rhythm changing process of step S1102.

[0227] First, in Figure 10A Shown in Figure 9Detailed example of the initialization process of step S901 in the embodiment, the CPU 201 performs the initialization process of TickTime. In this embodiment, the progress of the loop performance is performed in units of the value of the TickTime variable stored in the RAM 203 (hereinafter, the value of the variable is referred to as "TickTime", which is the same as the variable name). Figure 2 In the ROM 202 of the memory, the value of the TimeDivision constant (the value of this variable is hereinafter referred to as "TimeDivision", which is the same as the variable name) is preset, that is, the resolution of the quarter note. For example, when this value is 480, the duration of the quarter note is 480×TickTime. Note that the value of TimeDivision can also be stored in the RAM 203, and the user can change it, for example, by Figure 1 The switch on the first switch panel 102. In addition, as stored in Figure 2 The variables in RAM 203 of the present invention, TickTime, count: the value of the pointer variable used to determine that the user has performed a performance of a phrase (the value of the PhrasePointer value to be described later; the value of this variable is hereinafter referred to as "PhasePointer", the same as the variable name); the value of the pointer value used to record the loop performance (the value of the RecPointer value to be described later; the value of this variable is hereinafter referred to as "RecPointer", the same as the variable name); and the value of the pointer variable used for loop reproduction (the value of the PlayPointer value to be described later; the value of this variable is hereinafter referred to as "PlayPointer", the same as the variable name). The number of seconds that 1 TickTime actually corresponds to depends on the tempo specified for the song data. If the value set for the Tempo variable on RAM 203 according to the user setting is Tempo [beat / min], the number of seconds corresponding to 1 TickTime is calculated by the following equation.

[0228] TickTime[sec]=60 / Tempo / TimeDivision(10)

[0229] Therefore, in Figure 10A In the initialization process illustrated in the flowchart of , the CPU 201 first calculates TickTime [sec] by the calculation process corresponding to the above equation (10) and stores it in the variable of the same name on the RAM 203 (step S1001). Note that as the value of Tempo set for the variable Tempo, a value from 100 to 100 can be set in the initial state. Figure 2Alternatively, the variable Tempo may be stored in a nonvolatile memory so that the Tempo value at power-off is restored when the power of the electronic keyboard instrument 100 is turned on again.

[0230] Then, the CPU 201 calculates the TickTime [sec] calculated in step S1001 for the Figure 2 The timer 210 of the CPU 201 sets a timer interrupt (step S1002). As a result, each time TickTime [sec] elapses in the timer 210, an interrupt for the phrase progress of the sound synthesis LSI 205 and the loop recording / playback progress of the looper LSI 220 (hereinafter referred to as a "TickTime interrupt") occurs with respect to the CPU 201. Therefore, in the TickTime interrupt processing (described later) executed by the CPU 201 based on the TickTime interrupt, Figure 12 ), a control process is executed in which a phrase is determined every 1 TickTime and loop recording / playback is performed.

[0231] Subsequently, the CPU 201 performs various initialization processes, such as Figure 2 Initialization of RAM 203 (step S1003). Thereafter, CPU 201 ends Figure 10A The flowchart illustrates the Figure 9 Initialization processing of step S901.

[0232] Will be described later Figure 10B Flowchart of the process. Figure 11 It shows Figure 9 Flowchart of a detailed example of the switching process of step S902.

[0233] The CPU 201 first determines whether the rhythm change switch in the first switch panel 102 changes the rhythm of the phrase progress and loop recording / playback progress (step S1101). When the determination is "yes", the CPU 201 performs the rhythm change processing (step S1102). Figure 10B This processing is described in detail. When the determination in step S1101 is NO, the CPU 201 skips the processing of step S1102.

[0234] Then, the CPU 201 passes Figure 2 The key scanner 206 determines whether the user steps on the Figure 1 The pedal 105 is used to loop recording / playback in the looper LSI 220 (step S1103). When the determination is "yes", the CPU 201 executes the pedal control process (step S1104). Figure 14This processing is described in detail. When the determination in step S1103 is NO, the CPU 201 skips the processing of step S1104.

[0235] Finally, the CPU 201 executes other switching processing corresponding to the selection of the sound module tone of the electronic keyboard instrument 100. Figure 1 The second switch panel 103 is used to execute the musical instrument whose sound is to be synthesized in the sound synthesis LSI 205 (step S1105). The CPU 201 stores the sound module pitch and the musical instrument to be synthesized as a variable on the RAM 203 (not shown) (step S1104). Thereafter, the CPU 201 ends Figure 11 The flowchart illustrates the Figure 9 The switching process of step S902.

[0236] Figure 10B It shows Figure 11 Flowchart of a detailed example of the tempo change process of step S1102 in FIG. As described above, when the tempo value changes, TickTime [sec] also changes. Figure 10B In the flowchart of , the CPU 201 executes control processing related to the change of TickTime [sec].

[0237] First, something like Figure 9 The initialization process of step S901 is performed Figure 10A In the case of step S1001, the CPU 201 calculates TickTime [sec] by the calculation process corresponding to the above equation (10) (step S1011). Note that for the tempo value Tempo, it is assumed that Figure 1 The value of the rhythm change switch in the first switch panel 102 after the change is stored in the RAM 203 or the like.

[0238] Then, similar to Figure 9 The initialization process of step S901 is performed Figure 10A In the case of step S1002, the CPU 201 calculates TickTime[sec] in step S1011. Figure 2 The timer 210 sets the timer interrupt (step S1012). Thereafter, the CPU 201 ends Figure 10B The flowchart illustrates the Figure 11 The rhythm change processing of step S1102 is performed.

[0239] Figure 12 It is shown based on Figure 2 The TickTime interrupt that occurs every TickTime[sec] in the timer 210 (refer to Figure 10A Step S1002 or Figure 10B Flowchart of a detailed example of TickTime interrupt processing performed in step S1012).

[0240] First, the CPU 201 determines whether the value of a variable RecStart (hereinafter, the value of this variable is referred to as "RecStart" as the same as the variable name) on the RAM 203 is 1, that is, whether it indicates loop recording progress (step S1201).

[0241] When determining that the loop recording progress is not instructed (NO determination in step S1201 ), the CPU 201 proceeds to the processing of step S1206 without executing the processing of controlling the loop recording progress in steps S1202 to S1205 .

[0242] When it is determined that the loop recording progress is indicated (the determination of step S1201 is "YES"), in response to the TickTime interrupt advancing by 1 unit, the CPU 201 increases the value of the variable RecPointer on the RAM 203 (hereinafter, the value of this variable is referred to as "RecPointer", the same as the variable name) by 1, which is used to control the time progress in the TickTime unit in the loop segment so as to record the progress in the TickTime unit. Figure 3 In addition, the CPU 201 increases the value of the variable PhrasePointer (hereinafter referred to as "PhrasePointer" and having the same name as the variable) in the RAM 203 by 1 to control the time progress in the phrase segment in units of TickTime (step S1202).

[0243] Subsequently, the CPU 201 determines whether the value of PhrasePointer becomes the same as the value defined by TimeDivision×Beat (step S1203). As described above, the value of TimeDivision is based on TickTime, that is, the number of TickTimes corresponding to a quarter note. In addition, Beat is a value on RAM 203 (hereinafter, the value of this variable is referred to as "Beat", which is the same as the variable name), which stores a value indicating how many beats there are (how many quarter notes are included in one measure (one phrase)) for a piece of music that the user is about to perform. For example, if it is 4 beats, Beat=4, and if it is 3 beats, Beat=3. The value of Beat can be set by the user, for example, by Figure 1The switch on the first switch panel 102 is turned on. Therefore, the value of TimeDivision×Beat corresponds to the TickTime of one measure of the music currently being played. Specifically, in step S1203, the CPU 201 determines whether the value of PhrasePointer becomes the same as the TickTime of one measure (one phrase) defined by TimeDivision×Beat.

[0244] When determining that the value of PhrasePointer reaches TickTime of one measure (YES in step S1203), the CPU 201 Figure 1 The pitch sequence of each key pressed on the keyboard 101 (by Figure 9 The keyboard processing of step S903 stores it in the phrase buffer on the RAM 203) together with the data on the musical instrument synthesized by the sound specified in advance by the user during the segment with a TickTime of one measure (one phrase) (see Figure 9 In the switching process of step S902 Figure 11 Description of other switching processing of step S1105) is sent to Figure 2 The sound synthesis LSI205, as Figure 6 Then, the CPU 201 calculates the pitch data 215 of a phrase described in Figures 6 to 8 The statistical sound synthesis process described synthesizes the inference tone data 217 for one phrase and instructs output thereof to the looper LSI 220 (step S1204).

[0245] Thereafter, the CPU 201 resets the value of PhrasePointer to 0 (step S1205 ).

[0246] When determining that the value of PhrasePointer has not reached the TickTime of one measure (NO determination in step S1203 ), the CPU 201 executes only increment processing of RecPointer and PhrasePointer in step S1202 without executing the processing of steps S1204 and S1205 and shifts to the processing of step S1206 .

[0247] Subsequently, the CPU 201 determines whether the value of the variable PlayStart on the RAM 203 (hereinafter, the value of this variable will be referred to as "PlayStart", the same as the variable name) is 1, that is, indicates loop playback progress (step S1206).

[0248] When it is determined that the loop playback progress is indicated (the determination of step S1206 is "YES"), in response to the TickTime interrupt control by one unit, the CPU 201 increases the value of the variable PlayPointer on the RAM 203 (hereinafter, the value of this variable is referred to as "PlayPointer", the same as the variable name) by 1 for controlling Figure 3 The loop segment played back in a loop in the first loop storage area 301 or the second loop storage area 302 .

[0249] When determining that the loop playback progress is not indicated (NO determination in step S1206), the CPU 201 ends the Figure 12 The TickTime interrupt processing shown in the flowchart of , without executing the increment processing of PlayPointer in step S1207 and returning to execute Figure 9 Any one process of the main flowchart.

[0250] Then, reference will be made to Figures 13 to 15 Flowchart and Figure 16 and Figure 17 The operation is described in detail Figure 2 In the switching process of step S902 Figure 11 The pedal control process of step S1104 and Figure 2 The circulator LSI 220 is based on Figure 9 The looper control processing of step S905 implements the loop recording / playback processing from mode 0 to mode 3.

[0251] exist Figure 16 and Figure 17 In the operating instructions, Figure 16 t0 to Figure 17 The t22 in the figure represents a measure interval based on TickTime = a phrase interval = TimeDivision×Beat (reference Figure 12 In the following description, it is assumed that a description such as "time t0" means that the time is based on TickTime. In addition, in the following description, it is assumed that "bar" and "phrase" are used interchangeably and are unified as "bar".

[0252] Figure 13 It shows Figure 9 In the switching process of step S902 Figure 11 Flowchart showing details of the pedal control process at step S1104. First, it is assumed that immediately after the power of the electronic keyboard instrument 100 is turned on, for example, Figure 9 In step S901 Figure 10A In the miscellaneous initialization process of step S1003, Figure 13The value of the variable Mode on the RAM 203 shown (hereinafter referred to as "Mode", the same as the variable name), the value of the variable PrevMode (hereinafter referred to as "PrevMode", the same as the variable name) and the value of RecStart are all reset to 0.

[0253] exist Figure 13 In the pedal control process, Figure 11 In step S1103, Figure 2 After the key scanner 206 detects an operation on the pedal 105, the CPU 201 first detects the type of pedal operation (step S1301).

[0254] When it is determined in step S1301 that the user has stepped on the pedal 105 once with the foot or the like, the CPU 201 further determines the value of the current Mode.

[0255] When determining that the value of the current mode is 0, that is, when loop recording / playback is not being performed, the CPU 201 executes Figure 13 The series of processing from steps S1303 to S1308 in FIG. 1 is for switching from mode 0 to mode 1. Figure 16 The "STEP ON PEDAL FROM mode 0" state at time t0.

[0256] The CPU 201 first sets the value of the Mode variable to 1, indicating mode 1, and also sets the value of Mode 0 corresponding to the previous mode 0 to the PrevMode variable (step S1303).

[0257] Then, the CPU 201 sets the value of the RecStart variable to 1 to start loop recording in mode 1 (step S1304 ).

[0258] The CPU 201 then stores an indication of a variable RecArea (hereinafter, the value of this variable is referred to as "RecArea" which is the same as the variable name) on the RAM 203. Figure 3 The value Area1 of the first loop storage area 301 is used to represent Figure 3 The loop storage area where loop recording is performed on the circulator LSI 220 (step S1305).

[0259] Then, the CPU 201 sets a value of -TimeDivision×Beat for the variable RecPointer indicating the storage address in the TickTime unit of the loop recording, and also sets a value of 0 for the variable PhrasePointer (step S1306). Figure 12As described in step S1203, TimeDivision×Beat represents the tick time of one measure. Since the RecPointer value 0 is the start address for storage, the value -TimeDivision×Beat represents the time from the previous measure to the start of storage. This is because there is a one-measure delay until the RecPointer value returns to 0 during the TickTime interrupt processing. After the user starts a one-measure loop performance, there is a one-measure delay until the sound synthesis LSI 205 outputs the inferred musical tone data 217 corresponding to the loop performance.

[0260] The CPU 201 then sets the value of the PlayStart variable to 0 because loop playback is not performed in mode 1 (step S1307).

[0261] In addition, the CPU 201 sets the value of the LoopEnd variable indicating the end of loop recording on the RAM 203 (hereinafter, the value of this variable is referred to as "LoopEnd", which is the same as the variable name) to the value stored in the Figure 2 A sufficiently large number Max in the ROM 202 is read (step S1308).

[0262] By Figure 9 In the switching process of step S902 Figure 11 In the pedal control process of step S1104 Figure 13 After the user steps on the pedal 105 once in mode 0 to set mode 1, the CPU 201 performs a series of processing from step S1303 to step S1308 in the above-described manner. Figure 9 Step S905 Figure 14 The following processing is performed in the circulator control processing of .

[0263] exist Figure 14 In the case where the state transitions from mode 0 to mode 1, the CPU 201 first determines whether the RecStart value is 1 (step S1401). In the case where the state transitions from mode 0 to mode 1, the determination in step S1401 is "yes" because Figure 13 RecStart=1 has been set in step S1304.

[0264] When the determination in step S1401 is "YES", the CPU 201 determines whether the RecPointer value becomes equal to or greater than the LoopEnd value (step S1402). Figure 13 If a sufficiently large value is stored for the LoopEnd variable in step S1308, the determination in step S1402 is initially "No".

[0265] When the determination in step S1402 is NO, the CPU 201 determines whether loop recording is currently stopped and the RecPointer value becomes equal to or greater than 0 (step S1410 ).

[0266] exist Figure 13 In step S1306, for the RecPointer value, a negative value of one measure is stored, i.e. -TimeDivision×Beat. Figure 12 After the TickTime interrupt processing is changed to 1, every time the TickTime passes, the RecPointer value increases by 1 from -TimeDivision×Beat (see Figure 12 Until a bar of TickTime has passed and the RecPointer value becomes 0, Figure 14 The determination in step S1410 will continue to be "No", and Figure 15 The determination of whether the PlayStart value is 1 in the subsequent step S1412 also continues to be "No" (refer to Figure 13 Step S1307). Therefore, in Figure 14 Basically nothing is done in the looper control process, only time passes.

[0267] When the TickTime of one measure has passed and the RecPointer value has become 0, the CPU 201 makes Figure 2 Circulator LSI 220 Figure 3 The loop recording unit 303 starts from the area corresponding to the value Area1 indicated by the variable RecArea. Figure 3 The loop recording operation starts from the starting address 0 indicated by the variable RecPointer of the first loop storage area 301 (step S1411).

[0268] exist Figure 16 In the operation instructions, at time t0, as the user steps on the pedal 105, the state changes from mode 0 to mode 1. Figure 1 The key that specifies the pitch sequence of each measure on the keyboard 101 is pressed to start the loop performance, such as Figure 16 As shown in (a) at time t0. Figure 16 In the loop performance after time t0 shown in (a), the key designation is sent from the keyboard 101 to the sound module LSI 204 via the key scanner 206 and the CPU 201, so that the corresponding musical sound output data 218 is output from the sound module LSI 204 and the corresponding musical sound is emitted from the speaker (not shown) through the mixer 213, the digital-to-analog converter 211 and the amplifier 214.

[0269] Thereafter, as time passes, at time t1 after a TickTime of one measure has passed since time t0, the inference tone data 217 corresponding to the first measure starts to be output from the sound synthesis LSI 205. Figure 16 In synchronization with this, the loop recording of the inference musical tone data 217 after measure 1, which is sequentially output from the sound synthesis LSI 205 to the first loop storage area 301 (Area1), starts after time t1 by the processing of step S1411, as shown in FIG. Figure 16 At this time, for the performance input from measure 1 to measure 4, for example, Figure 16 The timing in (a) of FIG. 1 shows that the output timing and loop recording timing of the inference tone data 217 from the sound synthesis LSI 205 to the first loop storage area 301 are delayed by one measure, as shown in FIG. Figure 16 (b) and (c) of FIG. This is due to the limitation that the sound synthesis LSI 205 outputs the inferred tone data 217 with a one-measure delay relative to the note sequence of each measure input as the pitch data 215. Since the inferred tone data 217 delayed by one measure relative to the user's key performance is input to the looper LSI 220 but not output to the mixer 213, the corresponding sound generation is not performed.

[0270] The loop recording from section 1 is started at time t1 until the value of RecPointer reaches LoopEnd described later. The CPU 201 Figure 14 In the circulator control process of the repeat step S1401 "yes" judgment -> step S1402 "no" judgment -> step S1410 "no" judgment control. Thus, at time t1 Figure 14 In step S1411, Figure 2 Circulator LSI 220 Figure 3 The loop recording unit 303 continues the loop recording from RecPointer=0 (start address) to the area indicated by RecArea=Area1. Figure 3 The loop recording unit 303 sequentially loop records the data from Figure 16 (b) from time t1 to Figure 16 The time t5 described later shown in (c) is from Figure 2 The inferred musical tone data 217 of measures 1 to 4 is output from the sound synthesis LSI 205.

[0271] Then suppose the user Figure 16 During the loop performance shown in (a) of FIG, pedal 105 is stepped on at time t4, at the end of measure 4. As a result, Figure 9The switching process in step S902 corresponds to Figure 11 Step S1104 Figure 13 In the pedal control process, it is determined in steps S1301 and S1302 that the current mode is "step once" of Mode=1, so that the CPU 201 executes a series of processes from steps S1309 to S1314 to shift from Mode 1 to Mode 2.

[0272] The CPU 201 first sets the value of the Mode variable to 2, indicating mode 2, and also sets the value of Mode 1 corresponding to the previous mode 1 to the variable PrevMode (step S1309).

[0273] Then, the CPU 201 sets a value obtained by adding the value TimeDivision×Beat representing the TickTime of one measure to the current RecPointer value for the LoopEnd variable indicating the end of the loop recording (step S1310). Figure 16 In the example, the RecPointer value indicates time t4. Figure 16 As shown in (b) and (c) of FIG. 2 , as the output timing and loop recording timing of the inferred musical tone data 217, time t4 represents the end of measure 3 and is relative to Figure 16 The end of measure 4 in (a) is delayed by one measure. Therefore, to continue loop recording until time t5, which is delayed by one measure relative to the current RecPointer value, and complete recording to the end of measure 4, the value obtained by adding the value TimeDivision×Beat representing the TickTime of one measure to the previous RecPointer value is set to the LoopEnd variable, which indicates the end of loop recording. In other words, the LoopEnd value has a TickTime of four measures.

[0274] Subsequently, the CPU 201 sets the value of the variable PlayStart to 1 in order to enable loop playback by mode 2 for overdubbing (step S1311).

[0275] In addition, the CPU 201 sets the start address 0 for the PlayPointer variable indicating the address of the loop playback based on TickTime (step S1312).

[0276] Furthermore, for the variable PlayArea, the CPU 201 sets a value Area1 indicating the first loop storage area 301 that has been performed so far in a loop. The variable PlayArea indicates Figure 3 The loop storage areas in the first loop storage area 301 and the second loop storage area 302 shown are used for loop playback (step S1313).

[0277] Then, the CPU 201 makes Figure 2 Circulator LSI 220 Figure 3 The loop playback unit 304 starts from the value Area1 indicated by the variable PlayArea. Figure 3 The loop playback operation starts from the starting address 0 indicated by the variable PlayPointer in the first loop storage area 301 (step S1314).

[0278] exist Figure 16 In the operation instructions, at time t4, the user steps on the pedal 105, causing the state to change from mode 1 to mode 2. Figure 16 As shown at time t4 of (d), the user starts playing the key sequence of the pitch of each measure again from measure 1, and overlaps the key sequence from measure 1. Figure 2 The loop playback sound 310 starting from the start address (first measure) of the first loop storage area 301 for which loop recording has been performed so far in the looper LSI 220, and the loop playback sound 310 being emitted as the loop playback inferred tone data 222 from the speaker (not shown) via the D / A converter 211 and from the mixer 213 to the amplifier 214, as shown Figure 16 As shown in time t4 in (e). Figure 16 In the loop performance shown in (d) after time t4, the key designation is transmitted from keyboard 101 via key scanner 206 and CPU 201 to sound module LSI 204, resulting in the output of corresponding musical tone output data 218 from sound module LSI 204. The corresponding musical tone is then emitted from a speaker (not shown) via mixer 213, D / A converter 211, and amplifier 214. Simultaneously, as described above, loop playback inference musical tone data 222 output from circulator LSI 220 is also mixed with musical tone output data 218 by the user's loop performance in mixer 213 and then emitted. In this way, the user can execute the loop performance by pressing a key on keyboard 101 while listening to the loop playback inference musical tone data 222 recorded immediately before from circulator LSI 220.

[0279] Note that, as mentioned above Figure 13 As described in step 1310 of the embodiment, the variable LoopEnd is set to a value corresponding to time t5. Figure 9 Step S905 Figure 14 In the circulator control process, after time t4, until the Figure 12When the value of RecPointer, which is sequentially incremented every TickTime, reaches time t5 in step S1202 of the TickTime interrupt processing, the determination in step S1401 becomes "Yes", the determination in step S1402 becomes "No", and the determination in step S1410 becomes "No", so that the CPU 201 continues to input the inference tone data 217 of measure 4 from the sound synthesis LSI 205 to the looper LSI 220, as shown in FIG. Figure 16 As shown in (b), and as Figure 16 As shown in (c), the inference tone data 217 of measure 4 are loop-recorded (starting at step S1411) in the first loop storage area 301 (Area1).

[0280] Afterwards, when Figure 12 When the value of RecPointer, which is sequentially incremented every TickTime in step S1202 of the TickTime interrupt processing, reaches time t5, the determination in step S1401 becomes "Yes" and the determination in step S1402 also becomes "Yes". In addition, since the current mode is Mode = 1, the determination in step S1403 becomes "No". As a result, the CPU 201 first determines whether the value set for the variable RecArea is a value indicating Figure 3 The first loop storage area 301 of Area1, that is, whether the value is Figure 3 The second loop stores Area2 of the area 302 (step S1404). When the determination in step S1404 is "Yes" (RecArea=Area1), the CPU 201 changes the value of the variable RecArea to Area2 (step S1405). On the other hand, when the determination in step S1404 is "No" (RecArea≠Area1, i.e., RecArea=Area2), the CPU 201 changes the value of the variable RecArea to Area1 (step S1406). Figure 16 In the operation example, at time t5, the value of RecArea is Area1, such as Figure 16 As shown in (c), and Figure 3 The first loop storage area 301 is set as the target storage area for loop recording. However, after time t5, the value of RecArea becomes Area2, as shown in FIG. Figure 16 As shown in (h), Figure 3 The second loop storage area 302 newly becomes the target storage area for loop recording.

[0281] Thereafter, the CPU 201 sets the value of RecPointer to the start address 0 (step S1407 ).

[0282] Then, the CPU 201 makes Figure 2 Circulator LSI 220 Figure 3 The loop recording unit 303 starts the loop recording operation from the starting address 0 indicated by the variable RecPointer to the value corresponding to Area2 indicated by the variable RecArea. Figure 3 In the second loop storage area 302 (step S1408). Figure 16 In the operation example, Figure 16 In (h), this corresponds to time t5 and thereafter.

[0283] Then, in Figure 14 After step S1408, Figure 15 In step S1412, the CPU 201 determines whether the value of the variable PlayStart is 1. Figure 16 At time t4 in step S1311, since the value of the variable PlayStart is set to 1, the determination in step S1412 becomes "Yes".

[0284] Subsequently, the CPU 201 determines whether the value of the variable PlayPointer becomes equal to or greater than the value of the variable LoopEnd (step S1413). However, at time t4, since the value of the variable PlayPointer is still 0 (refer to Figure 13 1312), so the determination in step S1413 becomes "No". As a result, the CPU 201 ends Figure 14 and Figure 15 The flowchart shown Figure 9 Circulator control processing of step S905.

[0285] Like this, after the value of PlayPointer reaches LoopEnd, CPU 201 starts to play back in a loop from bar 1 at time t4. Figure 15 Repeat the circulator control process of "Yes" judgment of step S1412-> "No" judgment of step S1413 in the circulator control. Figure 16 (e) time t4 in Figure 13 In step S1314, Figure 2 The circulator LSI220 Figure 3 The loop playback unit 304 continues from the area indicated by PlayArea=Area1 Figure 3 The loop playback unit 304 starts from PlayPointer=0 (starting address) in the first loop storage area 301. Figure 16 The loop playback sound 310 is played back sequentially from measure 1 to measure 4 from time t4 to time t8 described in (e) and the loop playback sound 310 is emitted from the speaker.

[0286] At this time, according to the loop playback sound from measure 1 to measure 4 emitted from the speaker from time t4 to time t8 described later, the user presses Figure 1 The 101 key of the keyboard will continue the loop performance of each measure from measure 1 to measure 4, such as Figure 16 As shown in (d), the sound module LSI 204 issues musical tone output data 218 corresponding to the performance designation.

[0287] As a result, the pitch sequence and instrumental passages of each measure from measure 1 to measure 4 are Figure 12 The step S1204 is input to the sound synthesis LSI 205 on a measure basis. Figure 16 In the loop section from time t5 to time t9 described later in (f) of FIG. 1 , the sound synthesis LSI 205 outputs the inference tone data 217 with added rich musical expression to the Figure 2 The circulator LSI 220 has a delay of one bar.

[0288] On the other hand, from Figure 16 (e) The loop playback from time t4 to time t8 described later Figure 3 The loop playback sound 310 is composed of Figure 3 The phrase delay unit 305 delays one phrase so that the loop playback sound delay output 311 is input to the mixer 307 from time t5 to time t9 described later. Figure 16 As shown in (g). It can be seen that Figure 16 The inference tone data 217 input to the mixer 307 from time t5 to time t9 shown in (f) and Figure 16 As shown in (g), the loop playback sound delay output 311 input to the mixer 307 from time t5 to time t9 has the same timing from measure 1 to measure 4. Therefore, the mixer 307 mixes the inference tone data 217 and the loop playback sound delay output 311 and inputs them to Figure 3 As a result, the loop recording unit 303 sequentially adds the mixed audio data from measure 1 at time t5 to measure 4 at time t9 to the area represented by RecArea=Area2. Figure 3 302. Note that it is assumed that the operation of the phrase delay unit 305 is synchronized with the value of TimeDivision×Beat.

[0289] Here, assuming Figure 16 The loop playback sound 310 shown in (e) reaches the end of measure 4, which corresponds to the time t8 at which the loop segment ends. Figure 9 Step S905 Figure 14 and Figure 15 In the circulator control process of Figure 15 After the determination in step S1412 becomes "Yes", since the value of PlayPointer has reached the value of LoopEnd, the determination in step S1413 becomes "Yes". In addition, since the Mode value = 2 and the PrevMode value = 1 (reference Figure 13 , so the determination in step S1414 becomes "Yes". As a result, the CPU 201 first determines whether the value set for the variable PlayArea is a value indicating Figure 3 The first loop storage area 301 of Area1, that is, whether the value is Figure 3 Area2 of the second loop storage area 302 (step S1415). When the determination in step S1415 is "Yes" (PlayArea=Area1), the CPU 201 changes the value of the variable PlayArea to Area2 (step S1416). On the other hand, when the determination in step S1415 is "No" (PlayArea≠Area1, i.e., PlayArea=Area2), the CPU 201 changes the value of the variable PlayArea to Area1 (step S1417). Figure 16 In the operation example, at time t8, as Figure 16 As shown in (e), the value of PlayArea is Area1, and Figure 3 The first loop storage area 301 is set as the target storage area for loop playback. However, after time t8, the value of PlayArea becomes Area2, so that Figure 3 The second loop storage area 302 becomes the target storage area for loop reproduction. Figure 16 As shown in (j).

[0290] Thereafter, the CPU 201 sets the value of the variable PrevMode to 2 (step S1418 ).

[0291] Furthermore, the CPU 201 sets the value of the variable PlayPointer to the start address 0 (step S1419).

[0292] Then the CPU 201 makes Figure 2 Circulator LSI 220 Figure 3 The loop playback unit 304 starts from the value Area2 indicated by the variable PlayArea. Figure 3 The loop playback operation starts from the starting address 0 indicated by the variable PlayPointer in the second loop storage area 302 (step S1420). Figure 16 In the operation example, Figure 16 In (j), this corresponds to time t8 and thereafter.

[0293] on the other hand, Figure 16 The loop recording shown in (h) reaches the end of the loop segment at time t9 corresponding to the end of measure 4. In this case, Figure 9 Step S905 corresponds to Figure 14 and Figure 15 In the circulator control process of Figure 14 After the determination in step S1401 becomes "Yes", the determination in step S1402 becomes "Yes" because the value of RecPointer has reached the value of LoopEnd. In addition, since the Mode value = 2, the determination in step S1403 becomes "No". As a result, from step S1404 to step S1406, the CPU 201 executes the processing of exchanging the loop storage area in which the loop recording is performed. As a result, Figure 16 In the operation example, at time t9, as Figure 16 The value of RecArea shown in (h) is Area2, and Figure 3 The second loop storage area 302 is set as the target storage area for loop recording. However, after time t9, the value of RecArea becomes Area1, so Figure 3 The first loop storage area 301 becomes the target storage area for loop recording. Figure 16 As shown in (m).

[0294] Thereafter, the CPU 201 sets the value of RecPointer to 0 (step S1407 ).

[0295] Then the CPU 201 makes Figure 2 Circulator LSI 220 Figure 3 The loop recording unit 303 starts from the starting address 0 indicated by the variable RecPointer to the value Area1 corresponding to the value indicated by the variable RecArea. Figure 3 The loop recording operation in the first loop storage area 301 is performed (step S1408). Figure 16 In the operation example, Figure 16 In (m), this corresponds to time t9 and thereafter.

[0296] As described above, in Mode = 2, when loop playback and synchronized overdubbing reach the end of the loop section, the loop storage area in which loop recording has been performed so far and the loop storage area in which loop playback has been performed so far are swapped between the first loop storage area 301 and the second loop storage area 302, thereby continuing the overdubbing. This allows the user to generate the loop playback sound 310 of the loop section while sequentially overdubbing the inferred musical tone data 217 obtained based on the user's performance to the loop playback sound 310 of the loop section.

[0297] Assume that Figure 16 (i) and Figure 17 During the loop performance in (i), for example, in Mode=2, the user steps on the pedal 105 at any time at the end of measure 4, for example at time t12. As a result, Figure 9 The switching process in step S902 corresponds to Figure 11 Step S1104 Figure 13 In the pedal control process, it is determined in steps S1301 and S1302 that the current mode is Mode=2, which is "one step", so that the CPU 201 executes the operation from Figure 13 A series of processing from step S1315 to step S1321 is used for switching from mode 2 to mode 3.

[0298] The CPU 201 first sets the value of the Mode variable to 3 indicating mode 3, and also sets the value of Mode 2 corresponding to the previous mode 2 to the variable PrevMode (step S1315).

[0299] Then, the CPU 201 sets the start address 0 for the PlayPointer variable indicating the loop playback address based on the TickTime (step S1316).

[0300] Then, from Figure 13 In steps S1317 to S1319, the CPU 201 executes processing for exchanging the loop storage area in which loop playback is performed, similar to the process of FIG. Figure 15 A series of processing from step S1415 to step S1417 is performed. As a result, Figure 16 and Figure 17 In the operation example, at time t12, the value of PlayArea is Area2, such as Figure 17 As shown in (j), and Figure 3 The second loop storage area 302 is set as the target storage area for loop recording, but after time t12, the value of PlayArea becomes Area1, so that it has been the target storage area for overdubbing so far. Figure 3 The first loop storage area 301 becomes the target storage area for loop playback. Figure 17 As shown in (n).

[0301] Then, the CPU 201 sets a value obtained by adding the value TimeDivision×Beat representing the TickTime of one measure to the current RecPointer value for the LoopEnd variable indicating the end of the loop recording (step S1320). Figure 17 In the example shown in FIG, the RecPointer value indicates time t12. However, if Figure 17 As shown in (m), as the timing of loop recording, time t12 represents the end of section 3 and is relative to Figure 16 The end of measure 4 shown in (i) of FIG. 1 is delayed by one measure. Therefore, in order to continue loop recording until time t13 delayed by one measure relative to the current RecPointer value and complete recording until the end of measure 4, the value obtained by adding the value of TickTime indicating one measure, TimeDivision×Beat, to the current RecPointer value is set for the LoopEnd variable indicating the end of loop recording.

[0302] Then, the CPU 201 makes Figure 2 Circulator LSI 220 Figure 3 The loop playback unit 304 starts from the value Area1 indicated by the variable PlayArea. Figure 3 The loop playback operation starts from the starting address 0 indicated by the variable PlayPointer in the first loop storage area 301 (step S1321).

[0303] exist Figure 15 In step S1412, the CPU 201 determines whether the value of the variable PlayStart is 1. Figure 17 At time t12 in the operation example of MODE 2, since the value of the variable PlayStart has been set to 1 continuously from MODE 2, the determination in step S1412 becomes "YES".

[0304] Subsequently, the CPU 201 determines whether the value of the variable PlayPointer becomes equal to or greater than the value of the variable LoopEnd (step S1413). However, at time t12, since the value of the variable PlayPointer is 0 (refer to Figure 13 1316), so the determination in step S1413 becomes "No". As a result, the CPU 201 ends Figure 14 and Figure 15 The flowchart shown Figure 9 Circulator control processing of step S905.

[0305] In this way, Figure 17 As shown in (n), at time t12, loop playback starts from section 1 until the value of PlayPointer reaches LoopEnd, and the CPU 201 Figure 15 Repeat the circulator control process of "Yes" determination in step S1412->"No" determination in step S1413. Figure 17 Time t12 in (n) is at Figure 13 In step S1321, Figure 2 Circulator LSI 220 Figure 3 The loop playback unit 304 continues from the area indicated by PlayArea=Area1 Figure 3 The loop playback starts from PlayPointer=0 (starting address) of the first loop storage area 301. Then, the loop playback unit 304 starts from Figure 17 The loop playback sound 310 is played back sequentially from measure 1 to measure 4 from time t12 to time t16 described later in (n) and emitted from the speaker.

[0306] At this time, for example, the user can Figure 17 (n) described later in the time t12 to time t16 from the speaker to play back the loop sound from measure 1 to measure 4 to freely enjoy the performance, and by the key performance on the keyboard 101, in Figure 2 The sound module LSI 204 generates the tone output data 218. The tone output data 218 is mixed with the loop playback sound 310 in the mixer 213, and the tone output data 218 is then emitted from a speaker (not shown) through the D / A converter 211 and the amplifier 214.

[0307] As mentioned above Figure 13 As described in step 1310, the variable LoopEnd is set to a value corresponding to time t13. Figure 9 Step S905 Figure 14 In the circulator control process, after time t12 until Figure 12 When the value of RecPointer, which is sequentially incremented every TickTime, reaches time t13 in step S1202 of the TickTime interrupt processing, the determination in step S1401 becomes "Yes", the determination in step S1402 becomes "No", and the determination in step S1410 becomes "No", so that the CPU 201 continues inputting the inference tone data 217 of measure 4 from the sound synthesis LSI 205 to the looper LSI 220, as shown in FIG. Figure 16 As shown in (k), the input loop playback sound delay output 311, as shown in Figure 16As shown in (1), the inference tone data 217 of measure 4 is recorded in a loop in the first loop storage area 301 (Area1), as shown in FIG. Figure 16 As shown in (m).

[0308] Afterwards, when Figure 12 When the value of RecPointer, which is incremented by each TickTime, reaches time t13 in step S1202 of the TickTime interrupt processing, the determination in step S1401 becomes "Yes", and the determination in step S1402 also becomes "Yes". In addition, since the current mode is Mode=3, the determination in step S1403 becomes "Yes". As a result, the CPU 201 sets the value 0 to the RecStart variable for stopping loop recording (step S1409). Thereafter, the CPU 201 transfers to Figure 15 Step S1412.

[0309] Thus, the loop recording is ended in consideration of the delay of one phrase, and the loop recording is performed ("Yes" determination in step S1412->"Yes" determination in step S1413->"No" determination in step S1414->S1419->S1420). As a result, as Figure 17 As shown in (p), after loop playback is performed until time t20, the loop segment ends and the mode shifts to mode 3, in which loop playback starting from the first loop storage area 301 in which loop playback has been performed so far is performed from time t20, as shown in FIG. Figure 17 As shown in (t).

[0310] Assume that, although not Figure 17 (n) is shown in Figure 17 The loop playback sound 310 shown in (n) reaches the end of measure 4, which is time t16 corresponding to the end of the loop segment. Figure 17 The situation of (n) is different. If the user steps on the board, then Figure 15 After the determination in step S1412 becomes "Yes", the Figure 9 Step S905 Figure 14 and Figure 15 In the looper control process, since the value of PlayPointer has reached the value of LoopEnd, the determination in step S1413 becomes "Yes." Also, since the Mode value = 3, the determination in step S1414 becomes "No." As a result, the CPU 201 jumps to step S1419 and resets the value of PlayPointer to the start address 0 (step S1419).

[0311] Then the CPU 201 makes Figure 2 Circulator LSI 220 Figure 3 The loop playback unit 304 starts from the value Area1 indicated by the variable PlayArea. Figure 3 The loop playback operation is repeated from the starting address 0 indicated by the variable PlayPointer of the first loop storage area 301 (step S1420). Figure 16 Loop playback starting from the first loop storage area 301 from time t12 to time t16 in (n).

[0312] Then, suppose the user is at time t16, at the end of measure 4, e.g. Figure 17 The pedal 105 is depressed during the loop playback of the pattern 3 shown in (n). As a result, Figure 9 In the switching process of step S902 Figure 11 Step S1104 corresponds to Figure 13 In the pedal control process, it is determined in steps S1301 and S1302 that the current mode is Mode=3, which is "one step", so that the CPU 201 executes the operation from Figure 13 A series of processing from step S1322 to step S1323 is used for returning from mode 3 to mode 2.

[0313] The CPU 201 first sets the value of the Mode variable to 2 indicating mode 2, and also sets the value of Mode 3 corresponding to the previous mode 3 to the variable PrevMode (step S1322).

[0314] The CPU 201 then sets the value of the RecStart variable to 1 to start loop recording in mode 2 (step S1323 ).

[0315] The CPU 201 then determines whether the value set for the variable PlayArea is a value indicating Figure 3 The first loop storage area 301 of Area1, that is, whether the value is Figure 3 Area 2 of the second loop storage area 302 (step S1324). When the determination in step S1324 is "yes" (PlayArea = Area 1), that is, the loading source of the loop playback sound 310 played back so far in mode 3 is the first loop storage area 301, in order to use it as the loop playback sound 310 as it is, also in the next mode 2, and in order to set the loop storage area of ​​the loop recording to the second loop storage area 302, it is indicated Figure 3The value of the second loop storage area 302, Area2, is stored in the variable RecArea indicating the loop storage area for loop recording (step S1325). On the other hand, when the determination in step S1324 is "No" (PlayArea≠Area1), that is, the loading source of the loop playback sound 310 played back in mode 3 so far is the second loop storage area 302, in order to use it as the loop playback sound 310 as it is, also in the next mode 2, the loop storage area for loop recording is set to the first loop storage area 301, indicating Figure 3 The value Area1 of the first loop storage area 301 is stored in a variable RecArea indicating a loop storage area for loop recording (step S1326).

[0316] Then, the CPU 201 sets a value of -TimeDivision×Beat for a variable RecPointer indicating a storage address in TickTime units of loop recording and also sets a value of 0 for a variable PhrasePointer (step S1327 ). This processing is the same as that of step S1306 in Mode 1 .

[0317] After the transition from Mode 3 to Mode 2 Figure 17 The operation example at time t16 and thereafter in (o), (p), (q), (r), and (s) is similar to Figure 16 Operations at time t4 and thereafter in (d), (e), (f), (g) and (h).

[0318] Then, suppose that when the user has played to measure 2, for example Figure 17 During the overdubbing of mode 2 shown in (o), (p), (q), and (r) of FIG. 1 , the user holds the pedal 105 at time t19. As a result, Figure 9 In the switching process of step S902 Figure 11 Step S1104 corresponds to Figure 13 In the pedal control processing of , “HOLD” is determined in step S1301, so that the CPU 201 performs processing for transitioning to mode 3, in which only the last recording in overdubbing is canceled and the previous recording is played back.

[0319] That is, the CPU 201 first sets the value of the Mode variable to 3 (step S1327 ).

[0320] The CPU 201 then sets a value of 0 for the RecStart variable (step S1328).

[0321] Through the above processing, Figure 14In the looper control processing of step S1401, the determination in step S1401 becomes "No", so that the CPU 201 immediately (at time t19) stops the loop recording operation. In addition, after waiting for the PlayPointer value to reach LoopEnd, the CPU 201 leaves the loop storage area currently performing loop playback, returns the PlayPointer value to 0, and repeats the loop playback from the beginning ( Figure 15 "Yes" determination in step S1412 -> "Yes" determination in step S1413 -> "No" determination in step S1414 -> S1419 -> S1420). As a result, in the operation of Mode 2, Figure 17 The overdubbing shown by (q), (r) and (s) is canceled immediately after time t19, and Figure 17 The loop playback shown in (p) is performed until t20, the end of the loop segment, and the same loop playback sound 310 can be repeatedly played back from time t20 through mode 3. Thereafter, for example, the user can switch to mode 2 by stepping on the pedal 105 again and continue overdubbing.

[0322] Finally, the user can stop the cycling operation at any time by stepping on the pedal 105 twice. Figure 9 In the switching process of step S902 Figure 11 Step S1104 corresponds to Figure 13 In the pedal control process of , it is determined in step S1301 that "stepping on twice", so the CPU 201 first resets the Mode value to 0 (step S1329), and then resets all the RecStart values ​​and PlayStart values ​​to 0 (step S1330). As a result, Figure 14 and Figure 15 In the circulator control process of , the determinations in step S1401 and step S1412 all become "No," and the CPU 201 stops the loop control.

[0323] Usually, it is difficult to play a musical instrument while loop recording / playback operation. However, in the above embodiment, by playing a simple phrase and repeating the recording / playback operation, a loop recording / playback performance with rich musical expression based on the inference tone data 217 can be easily obtained.

[0324] As an embodiment using the inference tone data 217 output from the sound synthesis LSI 205, the embodiment in which the inference tone data 217 is used to Figure 2The looper LSI 220 is shown as an example of loop recording / playback. As another example using the inference music data 217 output from the sound synthesis LSI 205, an example is also conceivable in which the output of the inference music data 217 for a musical phrase is automatically recorded along with an automatic accompaniment or rhythm for enjoyment of an automatic performance. Thus, the user can enjoy an automatic performance that transforms a monotonous musical phrase performance into a musically expressive performance.

[0325] In addition, according to various embodiments of the present invention, by converting the performance phrases of the user's instrument sound into the performance phrases of the instrument sound by a professional performer, the converted instrument sound can be output, and a loop performance can be performed based on the output of the instrument sound.

[0326] According to the reference Figure 6 and Figure 7The first embodiment of the statistical sound synthesis process using the HMM acoustic model described above can smoothly reproduce the musical expression of a specific performer, a specific style, or the like, without splicing distortion, with smooth instrument sounds. Furthermore, by varying the training results 615 (model parameters), it is possible to adapt to other performers and express various musical phrases. Furthermore, since all model parameters in the HMM acoustic model can be automatically estimated from the training score dataset 611 and the training performance dataset 612, by using the HMM acoustic model to learn the characteristics of a specific performer, an instrument performance system can be automatically established for synthesizing sounds that reproduce these characteristics. Since the fundamental frequency and duration of the instrument sound output depend on the phrases in the musical score, the temporal structure of the pitch variation and rhythm over time can be easily determined from the score. However, the instrument sounds synthesized in this manner tend to be monotonous or mechanical, and are less appealing as instrumental sounds. In actual performances, in addition to standardized performances based on the musical score, stylistic characteristics of the performer or instrument can also be observed, such as variations in the pitch or on-time of notes, the accentuation of beats, and their temporal structure. According to the first embodiment of the statistical sound synthesis processing using the HMM acoustic model, because the changes over time in the spectrum data of the instrument performance sound and the pitch data as the sound source data can be modeled using context, it is possible to reproduce the instrument sound that is closer to the actual phrase performance. The HMM acoustic model used in the first embodiment of the statistical sound synthesis processing is a generative model that represents how the acoustic feature sequence of the instrument sound related to the vibration or resonance characteristics of the instrument, etc., changes over time during the phrase performance. In addition, according to the first embodiment of the statistical sound synthesis processing, by using the HMM acoustic model that takes into account the "deviation" context between the notes and the instrument sound, it is possible to achieve the synthesis of the instrument sound that can accurately reproduce the performance of the instrument that changes in a complex manner according to the performance technique of the performer. By combining the first embodiment of the statistical sound synthesis processing using this HMM acoustic model with a real-time phrase performance on, for example, the keyboard 100, it is possible to reproduce the phrase performance technique of the modeled performer (which has been impossible until now) to achieve a phrase performance as if the performer were actually playing in a keyboard performance on the electronic keyboard instrument 100.

[0327] When using reference Figure 6 and Figure 8In the second embodiment of the statistical sound synthesis process using a DNN acoustic model, a DNN replaces the HMM acoustic model, which relies on a decision tree-based context, as the representation of the relationship between the instrument sound feature sequence and the acoustic feature sequence in the first embodiment of the statistical sound synthesis process. This allows the relationship between the instrument sound feature sequence and the acoustic feature quantity sequence to be expressed using a complex nonlinear transformation function expressed using a decision tree. Furthermore, since the training data is classified according to the decision tree in the HMM acoustic model, which relies on the decision tree-based context, the amount of training data to be allocated to the HMM acoustic model depending on each context is limited. In contrast, in the DNN acoustic model, a single DNN learns from the entire training data, effectively utilizing the training data. Therefore, compared to the HMM acoustic model, the DNN acoustic model can more accurately predict the acoustic feature sequence, significantly improving the naturalness of the synthesized instrument sound. Furthermore, the DNN acoustic model can utilize the instrument sound feature sequence associated with the frame. Specifically, since the temporal correspondence between the acoustic feature sequence and the instrument sound feature sequence is predetermined in the DNN acoustic model, frame-related instrument sound features can be used, such as "the number of frames corresponding to the duration of the current note", "the position of the current frame in the note", etc., which are difficult to consider in the HMM acoustic model. In this way, by using frame-related instrument features, more accurate feature modeling can be achieved to improve the naturalness of the instrument sound to be synthesized. By combining the second embodiment of the statistical sound synthesis processing using such a DNN acoustic model with real-time performance on, for example, the keyboard 100, the performance of the instrument sound based on keyboard performance, etc. can be approximated to the performance technique of the modeled performer in a more natural way.

[0328] Although the present invention has been applied to the electronic keyboard musical instrument in the above-described embodiments, it can be applied to other electronic musical instruments such as electronic string instruments and electronic wind instruments.

[0329] In addition, the looper device itself can be an embodiment of an electronic musical instrument. In this case, by simply operating the user of the looper device to specify the pitch of a phrase and the embodiment of a specified musical instrument, a loop record / playback performance like a professional player playing it can be achieved.

[0330] Figure 6 The sound synthesis in the sound model unit 608 is not limited to the cepstrum sound synthesis. Another sound synthesis, such as LSP sound synthesis, may be used.

[0331] Although the statistical sound synthesis processing of the first embodiment using the HMM acoustic model and the second embodiment using the DNN acoustic model has been described as the above embodiments, the present invention is not limited thereto. Any type of statistical sound synthesis processing, such as an acoustic model combining HMM and DNN, may be employed.

[0332] Although the phrase consisting of the pitch sequence and the musical instrument is given in real time in the above embodiment, it may be given as a part of the automatic performance data.

[0333] This application is based on Japanese Patent Application No. 2019-096779 filed on May 23, 2019, the contents of which are incorporated herein by reference.

[0334] Reference Signs List

[0335] 100: Electronic keyboard instruments

[0336] 101: Keyboard (Performance Controls)

[0337] 102: First switch panel

[0338] 103: Second switch panel

[0339] 104: LCD

[0340] 105: Pedal (pedal operator)

[0341] 200: Control System

[0342] 201: CPU

[0343] 202: ROM

[0344] 203: RAM

[0345] 204: Sound Source LSI

[0346] 205: Sound Synthesis LSI

[0347] 206: Key Scanner

[0348] 208: LCD controller

[0349] 209: System Bus

[0350] 210: Timer

[0351] 211: Digital-to-Analog Converter

[0352] 213, 307: Mixer

[0353] 214: Amplifier

[0354] 215: Pitch data

[0355] 216: Sound production control data

[0356] 217: Musical instrument sound output data (inferenced musical sound data)

[0357] 218: Music output data

[0358] 219: Network interface

[0359] 220: Circulator LSI

[0360] 221: Beat data

[0361] 222: Loop playback of instrument sound output data

[0362] 301. Area1: First loop storage area

[0363] 302. Area2: Second loop storage area

[0364] 303: Loop recording unit

[0365] 304: Loop playback unit

[0366] 305: Phrase Delay Unit

[0367] 306: Beat extraction unit

[0368] 310: Loop playback sound

[0369] 311: Loop playback sound delay output

[0370] 600: Server

[0371] 601: Voice Training Unit

[0372] 602: Sound Synthesis Unit

[0373] 604: Training acoustic feature extraction unit

[0374] 605: Model training unit

[0375] 606: Trained acoustic model unit

[0376] 608: Sound Model Unit

[0377] 609: Oscillation generating unit

[0378] 610: Synthesis filter unit

[0379] 611: Training Music Dataset

[0380] 612: Training Instrument Sound Dataset

[0381] 614: Training Acoustic Feature Sequences

[0382] 615: Training Results

[0383] 617: Acoustic Feature Sequence

[0384] 618: Spectrum Data

[0385] 619: Sound source data

Claims

1. An electronic musical instrument comprising: show operator; as well as at least one processor, The at least one processor performs the following operations: Select the type of instrument; Based on pitch data associated with the performance operator operated by a user, acoustic feature data is digitally synthesized into sound data of a selected musical instrument and inferred musical tone data is output, the inferred musical tone data including inferred performance skills of a player different from the user, the inferred musical tone data being based on the acoustic feature data, the acoustic feature data being output by a trained acoustic model obtained by performing machine learning on: a training musical score dataset including training pitch data; and a training performance data set obtained by the performer playing the instrument, as played by the performer; and The inference musical tone data is not played by the user's operation of the performance operator.

2. The electronic musical instrument according to claim 1, wherein The inferred performance technique includes the articulation of legato, which is a notation in Western musical notation.

3. The electronic musical instrument according to claim 1 or 2, wherein: The at least one processor is configured to: acquiring first phrase data according to a first user operation, the first phrase data including a plurality of first notes from a first timing to a second timing; and First sound data corresponding to the first phrase and including the player's inference performance technique is repeatedly output.

4. The electronic musical instrument according to claim 3, wherein The at least one processor is configured to: acquiring, according to a second user operation, second phrase data including a plurality of second notes from a third timing to a fourth timing; and Second sound data corresponding to the second phrase and including the player's deductive performance technique is generated, the first sound data and the second sound data are mixed, and the mixed first sound data and second sound data are repeatedly output.

5. The electronic musical instrument according to claim 4, further comprising a pedal operator, wherein The at least one processor is configured to acquire the first phrase data or the second phrase data according to a user operation on the pedal operator.

6. The electronic musical instrument according to claim 4, wherein The at least one processor is configured to: shifting the note-on timing of the second note in the second phrase data to correspond to the note-on timing of the first note in the first sound data, the note-on timing not always being in agreement with the beat; and Second sound data shifted according to the note-on timing of the first note in the first sound data is generated.

7. The electronic musical instrument according to claim 6, wherein The at least one processor is configured to change the duration of the second notes in the second phrase data to correspond to the duration of the first notes in the first sound data, thereby changing the duration of at least one of the second notes in the second sound data.

8. A method performed by at least one processor of an electronic musical instrument, the method comprising the steps of: Select the type of instrument; digitally synthesizing acoustic feature data into sound data of a selected musical instrument based on pitch data associated with a performance operator operated by a user and outputting inferred musical tone data including inferred performance skills of a player different from the user, The inferred musical tone data is based on the acoustic feature data, which is output by a trained acoustic model obtained by performing machine learning on: a training musical score dataset including training pitch data; and a training performance data set obtained by the performer playing the instrument, as played by the performer; and The inferred musical tone data is not played by the user's operation on the performance operator.

9. The method according to claim 8, wherein The inferred performance technique includes the articulation of legato, which is a notation in Western musical notation.

10. The method according to claim 8 or 9, further comprising the steps of: acquiring, according to a first user operation, first phrase data, the first phrase data including a plurality of first notes from a first timing to a second timing; as well as The first sound data corresponding to the first phrase and including the player's inference performance technique is repeatedly output.

11. The method according to claim 10, further comprising the steps of: acquiring, according to a second user operation, second phrase data, the second phrase data including a plurality of second notes from a third timing to a fourth timing; as well as Second sound data corresponding to the second phrase and including the player's deductive performance technique is generated, the first sound data and the second sound data are mixed, and the mixed first sound data and second sound data are repeatedly output.

12. The method according to claim 11, further comprising the step of acquiring the first phrase data or the second phrase data in accordance with a user operation of a pedal operator.

13. The method according to claim 11, further comprising the steps of: shifting the note-on timing of the second note in the second phrase data to correspond to the note-on timing of the first note in the first sound data, the note-on timing not always being in agreement with the beat; and Second sound data shifted according to the note-on timing of the first note in the first sound data is generated.

14. The method according to claim 13, further comprising the steps of: The duration of the second notes in the second phrase data is changed to correspond to the duration of the first notes in the first sound data, thereby changing the duration of at least one of the second notes in the second sound data.

15. A computer-readable storage medium having stored thereon a program, the program being executable by at least one processor in an electronic musical instrument, the program causing the at least one processor to: Select the type of instrument; digitally synthesizing acoustic feature data into sound data of a selected musical instrument based on pitch data associated with a performance operator operated by a user and outputting inferred musical tone data including inferred performance skills of a player different from the user, The inferred musical tone data is based on the acoustic feature data, which is output by a trained acoustic model obtained by performing machine learning on: a training musical score dataset including training pitch data; and a training performance data set obtained by the performer playing the instrument, as played by the performer; and The inference musical tone data is not played by the user's operation of the performance operator.

Citation Information

Patent Citations

  • Automatic singing device

    JP1997050287A

  • Capacitor

    JP2019096779A

  • Automatic music continuation method and device

    US20020194984A1